Search PubMedSearch

SEARCH · Search PubMed

Results for “SNP Array”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Prenatal SNP-array chromosomal microarray analysis in 3,549 pregnancies: indication-specific yields and clinical implications.

BACKGROUND: SNP-based chromosomal microarray analysis (CMA) is widely used in invasive prenatal diagnosis, yet real-world performance across contemporary referral pathways, especially in the NIPT era, remains incompletely characterized. METHODS: We retrospectively analyzed 3,549 prenatal invasive samples tested by SNP array, and evaluated diagnostic yield overall and by referral indication and ultrasound phenotype. RESULTS: In total, we identified 398 pathogenic or likely pathogenic (P/LP) variants across 386 fetuses, resulting in an overall diagnostic yield of 10.9% (386/3,549). These findings comprised 223 aneuploidies and 175 pathogenic CNVs. In contrast, variants of uncertain significance (VOUS) were detected in 12.0% (426/3,549) of cases. Diagnostic yields were heavily stratified by indication: yields peaked in NIPT high-risk referrals (38.9%) and were intermediate in ultrasound-based cases (~ 11%), but dropped significantly in the advanced maternal age (AMA; 4.2%) and serum screening (~ 5-6%) groups. Conversely, VOUS rates remained remarkably stable across all referral categories. Sub-analysis of ultrasound abnormalities revealed that multisystem anomalies conferred the highest risk (27.3%), driven predominantly by aneuploidies; among soft markers, increased nuchal translucency (NT) emerged as the strongest predictor of chromosomal pathology. CONCLUSIONS: In our cohort, SNP-array identified clinically actionable findings in 10.9% of cases. NIPT enriched diagnostic yields, particularly for aneuploidies, and NT thickness was strongly associated with pathogenic findings. These results support an indication-based approach to genomic testing, with NIPT as a triage tool for aneuploidy and CMA for high-risk populations, while improving VOUS counseling.

Humans

High-Density SNP Genotyping Reveals High Population Connectivity and Limited Spatial Genetic Structure in Apodemus flavicollis and Apodemus sylvaticus.

High-density SNP arrays are increasingly used in ecological and evolutionary studies, yet their application in wild species remains challenging. In this study, we evaluated the performance of the Affymetrix Axiom Mouse HD array, originally developed for Mus musculus, in two wild small mammals, Apodemus flavicollis and Apodemus sylvaticus, with particular focus on genetic diversity and population connectivity across seven sampling sites within a fragmented landscape. A total of 96 individuals (43 A. flavicollis and 53 A. sylvaticus) were genotyped using a 616K SNP array. After quality control filtering for missingness and minor allele frequency, more than 160,000 high-quality autosomal SNPs were retained for each species. Despite being designed for a different species, the array effectively discriminated between A. flavicollis and A. sylvaticus, with principal component analysis clearly separating the two species. Levels of genetic diversity were comparable across sites, with mean observed heterozygosity around 0.33 and consistently negative F IS values, indicating a slight excess of heterozygotes. Population structure analyses revealed extremely weak spatial genetic differentiation. ADMIXTURE supported a single genetic cluster (K = 1) within each species, while analysis of molecular variance attributed more than 99% of genetic variation to within-individual components. Pairwise relationship analyses showed that related individuals were not confined to single sites but occurred across sampling locations, supporting ongoing gene flow even across the fragmented landscape. No significant isolation-by-distance pattern was detected. Overall, our results indicate high population connectivity and limited spatial genetic structuring in both species across the study area, consistent with the documented dispersal capacity of these species at the spatial scale investigated. Moreover, this study demonstrates that high-density SNP arrays can provide powerful genomic tools for investigating dispersal dynamics and population structure in closely related wildlife species under habitat fragmentation, where subtle genetic patterns may otherwise remain undetected.

Apodemus species

Scalable medium-density genotyping platforms for cultivar identification, pedigree authentication, marker-assisted and genomic selection, and other applications in strawberry.

A broad spectrum of high-density genotyping approaches, including single-nucleotide polymorphism (SNP) arrays, genotyping-by-sequencing, and whole-genome reduced-representation sequencing, have been shown to perform well in strawberry (Fragaria × ananassa), despite the inherent complexity of the octoploid genome. While these approaches are effective, their routine deployment in breeding programs can be constrained by cost, computational requirements, and workflow complexity. In parallel, many breeding programs continue to rely on locus-specific assays for marker-assisted selection, resulting in fragmented and inefficient genotyping strategies. Here, we describe medium-density amplicon-based genotyping platforms for strawberry designed to provide cost-effective, turnkey solutions that integrate markers used for marker-assisted selection with genome-wide markers suitable for genomic prediction in a single laboratory assay. These platforms were developed by targeting 1,650 or 4,811 target SNPs via amplicon sequencing, and are interoperable with existing high-density genotyping resources, including a widely used 50K SNP array, thereby facilitating data integration across platforms. We benchmarked their performance relative to the 50K SNP array across breeding-relevant applications, including identity and purity testing, pedigree authentication, marker-assisted selection, and genomic selection, and further evaluated the feasibility of genotype imputation to enhance genome-wide information content. Across analyses, the 1,650- and 4,811-amplicon platforms produced results comparable to higher-density platforms while substantially reducing genotyping cost and analytical overhead. This work demonstrates that targeted amplicon-based genotyping can support efficient, scalable, and integrated genome-informed breeding, enabling the routine application of both marker-assisted and genomic selection within strawberry breeding workflows. Open-source R workflows are provided to support streamlined analyses in breeding contexts.

Fragaria

Assessment of genomic prediction capabilities of transcriptome data in a barley multi-parent RIL population.

Low-cost and high-throughput RNA sequencing data for barley RILs achieved GP performance comparable to or better than traditional SNP array datasets when combined with parental whole-genome sequencing SNP data. The field of genomic selection (GS) is advancing rapidly on many fronts including the utilization of multi-omics datasets with the goal of increasing prediction ability and becoming an integral part of an increasing number of breeding programs ensuring future food security. In this study, we used RNA sequencing (RNA-Seq) data to perform genomic prediction (GP) on three related barley RIL populations. We investigated the potential of increasing prediction ability by combining genomic and transcriptomic datasets, adding whole-genome sequencing (WGS) SNP data, functional annotation-based filtering, and empirical quality filtering. Our RNA-Seq data were generated cost-efficiently using small-footprint plant cultivation, high-throughput RNA extraction, and Library preparation miniaturization. We also examined sequencing depth reduction as an additional cost-saving measure. We used fivefold cross-validation to evaluate the prediction ability of the gene expression dataset, the RNA-Seq SNP dataset, and the consensus SNP dataset between the RNA-Seq and parental WGS data, resulting in prediction abilities between 0.73 and 0.78. The consensus SNP dataset performed best, with five out of eight traits performing significantly better compared to a 50K SNP array, which served as a benchmark. The advantage of the consensus SNP dataset was most prominent in the inter-population predictions, in which the training and validation sets originated from different RIL sub-populations. We were therefore able to not only show that RNA-Seq data alone are able to predict various complex traits in barley using RILs, but also that the performance can be further increased with WGS data for which the public availability will steadily increase.

Hordeum

A vision of how low-coverage sequence data should contribute to genetic evaluation in the future.

Low-coverage sequencing refers to sequencing DNA of individuals to a low depth of coverage (e.g., 0.5X) and imputing that sequence to a genomic sequence based on reference haplotypes from individuals sequenced to a high depth of coverage (e.g., ≥10X). It has been proposed as an alternative to genotyping by Single-nucleotide polymorphisms (SNP) arrays. At least one commercial product based on it is available for agricultural species. Concerns limiting adoption in its current form are: 1) the cost of storing the huge volume of data it generates and 2) whether that additional data will result in improved accuracy of genetic evaluation. This work envisions future implementation of low-coverage sequencing to reduce storage costs and enhance genetic evaluations by leveraging the additional information in the full sequence of the pangenome to account for more genetic variation. We propose addressing the storage issue by representing genomic sequence of an individual in a pair of haplotype arrays with each element pointing to an enumerated haplotype of the sequence within one of approximately 50,000 defined genome segments. Assuming 60 million genomic variants, the infrastructure required to translate the identifier of any enumerated haplotype into its genomic sequence would require less than 10 gigabytes of binary storage. Each haplotype array element would require 2 bytes, so the marginal binary storage required to represent the genomic sequence of an individual would be about 200 kilobytes (KB), similar to the genotypes from a SNP array with 200,000 markers. This assumes no pedigree and no ambiguity of the imputation, though the latter is unrealistic. Strategies to minimize, and when necessary, to manage and efficiently represent ambiguity are proposed. The genomic sequence of an individual could be stored in about 1 KB (binary) if both parents have unambiguous sequences stored as described above. The proposed system for representing the pangenome includes algorithms for read mapping and imputation intended to leverage all known genetic variation in the target population. It is also designed to use sequencing reads generated for imputing the genomic sequence of new individuals to identify unrecognized mutations, crossovers, and structural variants, thus continuously improving the genome representation, especially if widespread use of low-coverage sequencing in livestock industries is realized. This could make improved genetic merit and management of livestock feasible without computational burden.

Animals

[Optical genome mapping analysis of a Chinese pedigree with a complex balanced translocation involving four chromosomes].

OBJECTIVE: To explore the genetic characteristics of a complex balanced translocation involving four non-homologous chromosomes in a Chinese pedigree using optical genomic mapping (OGM). METHODS: A woman with primary infertility and her family members who presented at the Prenatal Diagnosis Center of the Sixth Affiliated Hospital of Sun Yat-sen University in October 2021 were selected as study subjects. Comprehensive analysis and verification of chromosomal abnormalities were conducted through conventional G-band karyotyping analysis, single nucleotide polymorphism microarray (SNP array) and OGM. This study was approved by the Medical Ethics Committee of the hospital (Ethics No.: E2022210). RESULTS: G-band karyotyping analysis indicated that the proband, her father, and younger brother have all carried a complex translocation involving four chromosomes. SNP array analysis revealed a duplication of approximately 21.63 Mb in the 9p24.1-p21.1 region in the proband's younger brother, while no abnormality was detected in other family members. OGM confirmed that the complex balanced translocation has involved chromosomes 5, 8, 9, and 10. CONCLUSION: The proband has harbored a complex balanced translocation. OGM has demonstrated certain advantages in characterization of complex chromosomal structural abnormalities.

Humans

Pooled DNA genotyping on Affymetrix SNP genotyping arrays.

BACKGROUND: Genotyping technology has advanced such that genome-wide association studies of complex diseases based upon dense marker maps are now technically feasible. However, the cost of such projects remains high. Pooled DNA genotyping offers the possibility of applying the same technologies at a fraction of the cost, and there is some evidence that certain ultra-high throughput platforms also perform with an acceptable accuracy. However, thus far, this conclusion is based upon published data concerning only a small number of SNPs. RESULTS: In the current study we prepared DNA pools from the parents and from the offspring of 30 parent-child trios that have been extensively genotyped by the HapMap project. We analysed the two pools with Affymetrix 10 K Xba 142 2.0 Arrays. The availability of the HapMap data allowed us to validate the performance of 6843 SNPs for which we had both complete individual and pooled genotyping data. Pooled analyses averaged over 5-6 microarrays resulted in highly reproducible results. Moreover, the accuracy of estimating differences in allele frequency between pools using this ultra-high throughput system was comparable with previous reports of pooling based upon lower throughput platforms, with an average error for the predicted allelic frequencies differences between the two pools of 1.37% and with 95% of SNPs showing an error of < 3.2%. CONCLUSION: Genotyping thousands of SNPs with DNA pooling using Affymetrix microarrays produces highly accurate results and can be used for genome-wide association studies.

Alleles

Development and validation of a high-density 'Amahysnp' genotyping array in grain amaranth (Amaranthus hypochondriacus).

BACKGROUND: Grain amaranth has recently gained global attention as a promising crop alternative to traditional cereals due to its nutritional value and adaptability to various growing conditions. Although gene banks conserve extensive collections of amaranth germplasm, the genomic and phenotypic characterization of these resources is limited, which hinders their full utilization in breeding programs. A major challenge is the lack of high-throughput genotyping assays essential for comprehensive genomic characterization and trait mapping. High-density SNP arrays have become standard tools for genome-wide analysis across multiple loci, enabling molecular breeding across a range of crop species. RESULTS: In this study, we developed a 64&#xa0;K high-throughput SNP genotyping array named "AmahySNP", using Affymetrix&#xae; Axiom&#xae; technology. The array contains 64,069 high-density SNPs distributed across both genic (55.17%) and non-genic (44.83%) regions of the Amaranthus hypochondriacus genome. The genic region includes 8,879 genes, which consist of 4,830 single-copy genes and 4,049 multi-copy genes distributed across 16 scaffolds. These genes cover various functional regions, including exons (10.5%), introns (40.1%), 5'UTRs (1.6%), and 3'UTRs (2.9%), respectively. The AmahySNP array was effectively utilized for population structure analysis, genetic diversity studies, core development, and genome wide association studies (GWAS) in amaranth germplasm. A representative core set of 112 accessions was identified, which includes two released varieties (Annapurna and Suvarna) and 100 diverse accessions from 12 different regions, representing 12% of the total 917 accessions evaluated. Phylogenetic analysis revealed three major genetic clusters, independent of their geographical origins. GWAS conducted using 22,763 polymorphic SNPs from 540 genotypes identified 13 novel loci associated days to flowering (DTF) trait, seven of which were located within annotated genes. CONCLUSIONS: The AmahySNP 64&#xa0;K SNP chip a valuable genomic tool for amaranth research and breeding with a strong potential to accelerate its genetic improvement. It enables high-throughput genotyping for a wide range of applications, including GWAS and other genomic studies, and will significantly advance the exploration of natural genetic variations. Ultimately, this resource will empower amaranth breeders to develop improved amaranth cultivars with enhanced crop yield, resilience, and nutritional quality, contributing to global food security and sustainable agriculture.

Amaranthus

Inference of Genetic Structure and the Process of Population Formation in Nepalese Native Goats Using Uniparental and Genome-Wide Markers.

Nepal is a small, landlocked country with marked elevational variation from the Terai plains to the Himalayas. Here, four indigenous goat populations (Chyangra, Sinhal, Khari, and Terai) are raised at different elevations. This study aimed to clarify the genetic structure of these populations and how they are formed and propagated across the Himalayan region. We analyzed 136 Nepalese goats using mitochondrial (mt) DNA D-loop and sex-determining region Y (SRY) 3'-untranslated region (UTR) sequences, as well as 50 K SNP array data. The mtDNA haplogroups D (0.162) and G (0.03) were detected only in Chyangra, whereas haplogroup B was predominant in Sinhal (0.42), followed by Khari (0.260). Regarding SRY haplotypes, Y2B was detected in all populations, whereas Y1AB (0.42) was found only in Chyangra. Genome-wide SNP analysis showed that Chyangra was genetically related to Tibetan and Central Asian goats, while Terai resembled South Asian goats. Interestingly, Sinhal formed a distinct cluster, whereas Khari exhibited an admixed genetic structure. These findings suggest that Nepalese goats originate from at least three ancestral lineages and that an additional migration route may have existed through the southern Himalayas.

50K SNP

STK11 Mutations and Deletions Define an Aggressive Molecular Subgroup of Cervical Adenocarcinoma.

Cervical adenocarcinoma accounts for 15%-20% of cervical cancers and is associated with poorer survival and reduced response to screening and immunotherapy compared with squamous cell carcinoma (SCC). The genomic drivers underlying this molecular subgroup remain incompletely characterized. Whole-exome sequencing was performed on 302 invasive cervical cancers from Guatemala and Venezuela. Structural variation analysis was conducted using SNP-array and whole-genome sequencing data. Findings were replicated in more than 4600 additional cervical cancer samples from TCGA, AACR Project GENIE, MSKCC, and Caris datasets. TP53 mutations were more frequent in adenocarcinoma than SCC, particularly in HPV-negative tumors. STK11 alterations, including mutations and focal deletions, were significantly enriched in HPV-positive adenocarcinomas compared with SCC and affected 23% of adenocarcinomas overall. Whole-genome analyses identified recurrent focal deletions, inversions, chromosomal rearrangements, and breakage-fusion-bridge events involving chromosome 19p and STK11 that were not detected by exome sequencing alone. STK11 alterations were associated with younger age at diagnosis, poorer overall survival, and inferior outcomes following immune checkpoint inhibitor (ICI) therapy. STK11 alterations significantly co-occurred with YAP1 amplification but were largely mutually exclusive with PIK3CA mutation. Cervical adenocarcinomas also demonstrated significantly lower CD274 (PD-L1) expression than SCC. STK11 alterations define a distinct molecular subgroup of cervical adenocarcinoma characterized by structural disruption of chromosome 19p, younger age at onset, and poorer clinical outcomes. These findings have implications for molecular classification and future targeted therapeutic approaches in cervical cancer.

Humans

Cribriform tumors of the skin: CD38 expression distinguishes from histologic mimics in a multi-institutional series.

Cribriform tumor of the skin (formerly primary cutaneous cribriform carcinoma/primary cutaneous cribriform apocrine carcinoma) is a rare adnexal neoplasm of uncertain malignant potential. Recent studies have identified recurrent co-deletion of the long arm of chromosomes 6 and 9 and CD38 overexpression as potential diagnostic features, but their sensitivity and specificity remain incompletely defined. We identified 11 cribriform tumors with classic morphologic features, including several previously reported molecular characterized cases and additional unpublished cases, from multiple institutions. CD38 immunohistochemistry was performed on all tumors and compared with a panel of morphologic mimics, including adenoid cystic carcinoma (n&#x2009;=&#x2009;8), digital papillary adenocarcinoma (n&#x2009;=&#x2009;8), eccrine/apocrine adenomas (n&#x2009;=&#x2009;7), hidradenoma (n&#x2009;=&#x2009;2), and endocrine mucin-producing sweat gland carcinoma (n&#x2009;=&#x2009;3). SNP array analysis was available in nine cases. Patients had a median age of 47 years with a slight female predominance (64%). Tumors commonly involved the extremities and ranged from 0.3 to 2.0 cm. Histologically, all cases demonstrated a well-circumscribed dermal neoplasm of bland epithelial cells arranged in cribriform architecture with characteristic thread-like intraluminal bridging, without high-grade features. CD38 expression was identified in 9/11 cases (82%), typically with moderate-to-strong diffuse cytoplasmic staining. All morphologic mimics (n&#x2009;=&#x2009;28) were CD38 negative. Recurrent chromosomal deletions involving 6q and/or 9q were identified in 7/8 (88%) successfully tested cases. Two morphologically classic cribriform tumors lacked CD38 expression, including one with confirmed 6q/9q codeletion, while one CD38-positive case lacked detectable genomic copy number alterations. These findings further support cribriform tumors as a molecularly distinct, low-grade adnexal neoplasm characterized by recurrent 6q/9q deletions and frequent CD38 expression. CD38 appears to be a useful adjunctive diagnostic marker with high specificity among the selected mimics tested, although its sensitivity is incomplete, and absence of staining does not exclude the diagnosis.

Adnexal tumors

MetaGLIMPSE: Meta-imputation of low-coverage sequencing data for modern and ancient genomes.

The advent of efficient and accurate imputation for low-coverage sequencing offers an unbiased alternative to SNP array imputation, increasing the accuracy of rare variant imputation across all populations. Since imputation accuracy generally increases with larger reference panels and closer ancestry match between target and reference samples, leveraging imputation from multiple reference panels improves imputation accuracy; however, individual reference panel genotypes are often privacy protected. Meta-imputation bypasses individual-level data by combining single-panel imputed genotypes through estimating panel- and marker-specific weights. We present a meta-imputation method, MetaGLIMPSE, that combines estimates from multiple reference panels for low-coverage sequencing imputation. Across all our scenarios, for both modern and ancient DNA samples, MetaGLIMPSE consistently outperforms the best single-panel imputation for coverages of 0.1&#xd7;-8&#xd7; and across all minor-allele frequencies, equaling the combined panel imputation for some parameters. Finally, MetaGLIMPSE is computationally efficient, meta-imputing 500 whole genomes in 16% of the time of GLIMPSE2.

Humans

Genome-wide SNP-based genomic diversity and population structure analysis in alpaca populations from Europe and Peru.

This study aimed to analyze the genetic diversity and population structure of alpacas in Germany, Switzerland, and Austria (German-speaking regions, GSR) and to compare with that of the country of origin of the species (Peru). A total of 179 animals from GSR and 151 from Peru were genotyped with a species-specific 76k SNP array. The observed and expected heterozygosity was 0.305 and 0.311 for GSR and 0.310 and 0.312 for Peru. The mean FROH values were 0.029 for GSR and 0.023 for Peru. In general, results show that breeders in both analyzed regions efficiently maintain genetic diversity. Principal component analysis identified the GSR and Peru populations as separate from each other, but the relative proximity of both clusters indicates the shared genetic heritage. FST and XPEHH methods identified genomic regions under selection for traits such as coat color and adaptation. Genome-wide association studies comparing black and brown with white or gray alpacas identified associated genome regions containing the ASIP and KIT genes, respectively. The association of a recently identified keratin locus on chromosome 16 with differences in fleece type in alpacas was confirmed, while the putative causality of a TRPV3 variant was rejected.

Animals

The curious case of sporadic nematode susceptibility in "Tifguard" peanut (Arachis hypogaea): seed mixture or genetic instability?

The Runner-type peanut (Arachis hypogaea L.) cultivar "Tifguard" carries an introgressed chromosomal segment on chromosome A09 from A. cardenasii that confers resistance to root-knot nematode (RKN). Despite this, a proportion of "Tifguard" plants show RKN symptoms, which could plausibly be attributed to seed mixture or outcrossing. However, recent work has shown that cultivated peanut exhibits surprisingly frequent large-scale chromosomal instability (1% to 5%); suggesting that resistance loss could arise from spontaneous structural genomic change. To test these possibilities, we grew foundation seed in an RKN-infested field and collected symptomatic and asymptomatic plants. Lineages derived by single-seed descent were genotyped using the Axiom Arachis 48K SNP array v2 and whole-genome sequencing. Symptomatic lineages lacked the A. cardenasii introgression on chromosome A09 and instead carried the complete endogenous A. hypogaea A09 region at the expected dosage. There was no evidence of large-scale homoeologous exchange, deletion, or other genomic instability affecting this chromosome. Most susceptible plants were closely related to resistant "Tifguard" but lacked the A09 introgression, with a smaller proportion assignable to known nematode-susceptible cultivars, implicating seed mixture with a possible contribution from cross-pollination rather than genomic instability. Because resistance depends on a single major-effect segment, rare events have disproportionate phenotypic impact, placing high demands on genetic purity. For important traits conferred by major loci, marker-based testing across seed-increase stages could verify trait retention directly, and is increasingly practical as marker costs decline.

Arachis

Accurate identification of abnormal ploidy using an artificial intelligence model in preimplantation genetic testing.

STUDY QUESTION: Can ultra-low-coverage whole-genome sequencing (ulc-WGS) accurately identify abnormal ploidy during preimplantation genetic testing (PGT)? SUMMARY ANSWER: The artificial intelligence (AI)-based PGT-Plus model demonstrates high accuracy in ploidy detection, offering a cost-effective solution that enhances clinical utility of PGT. WHAT IS KNOWN ALREADY: The predominant PGT for aneuploidy can identify chromosomal aneuploidies but cannot determine ploidy status. Transferring embryos with ploidy abnormalities can result in miscarriage and molar pregnancy. On the other hand, in ART, fertilization is assessed by morphological pronuclear assessment at the zygote stage. However, it has a low specificity in the prediction of abnormal ploidy status and embryos deemed abnormally fertilized can yield healthy pregnancies. Accurately identified abnormal ploidy in PGT-A can resolve current limitations and expand the utility range of PGT-A. Several studies have identified ploidy abnormalities; however, they were mainly based on single-nucleotide polymorphism (SNP) arrays or needed to combine additional targeted-next-generation sequencing (NGS) information. Studies based on ulc-WGS remain scarce. STUDY DESIGN SIZE DURATION: The study consisted of two stages: methodology establishment and validation. An AI model, named PGT-Plus, was developed using 653 samples with known ploidy status, which was further validated using 792 different ploidy status samples. In the clinical application stage, the approach was used to analyse the ploidy status of 19&#x2009;103 normally fertilized PGT blastocysts and 140 single pronucleus (1PN)-derived blastocysts collected between May 2022 and December 2023. All blastocysts were tested using trophectoderm biopsy and NGS. PARTICIPANTS/MATERIALS SETTING METHODS: The methodology is based on the ulc-WGS data. First, based on samples with known ploidy status: the heterozygosity rate of high-frequency biallelic SNPs, the likelihood ratio (LLR) of alleles was calculated under different assumptions ('both parental homologs' [BPH] from a single parent, 'single parental homolog' [SPH] from each parent, disomy, and monosomy) by leveraging allele frequencies and linkage disequilibrium (LD) measured in the 1000 genomes project database. Twenty-three continuous candidate features derived from heterozygosity rates and LLRs of chromosomes or selected windows were included to establish the ploidy prediction AI model. Gini importance analysis and multicollinearity mitigation was performed for feature selection, then the performance of Random Forest (RF), Support Vector Machine (SVM), and Logistic Regression for modelling was compared. Subsequently, the parameter optimization was performed based on the RF model. Ploidy constitution concordance was evaluated in known ploidy status samples. The frequency of abnormal ploidy in normal fertilized PGT blastocysts and 1PN-derived blastocysts (including conventional IVF and ICSI) was evaluated. MAIN RESULTS AND THE ROLE OF CHANCE: Eleven features were collected for model architecture compared to SVM and Logistic Regression; RF achieved superior performance for ploidy detection. The AI model achieved an AUC of 1 for genome-wide-uniparental diploidy (GW-UPD), 1 for triploidy, and 0.99 for diploidy. For the 792 validation samples, 99.5% of samples were successfully detected using the AI model, and the model showed 100% accuracy for ploidy classification. In the clinical application stage, out of 19&#x2009;103 PGT samples, 19&#x2009;069 were successfully analysed using the model, with 110 (0.57%) identified as having abnormal ploidy embryos. Among these, 12.7% (14/110) were identified as GW-UPD, and 87.3% (96/110) were triploid. Among 5563 diploid blastocysts transferred, 3478 clinical pregnancies were achieved. Subsequent ploidy analysis was performed for 217 spontaneous abortion and 935 prenatal diagnostic samples, and no abnormal ploidy was identified. Furthermore, of the 140 1PN embryos tested, 40 (28.6%) exhibited GW-UPD, 3 (2.1%) exhibited triploidy, and 97 (69.3%) were determined to be biparental and normally fertilized. Among the 97 biparental embryos, 46 were diploid, 11 were mosaic, and 40 were aneuploid. In terms of the insemination pattern, the percentage of abnormal ploidy in ICSI was significantly higher than in conventional IVF (P&#x2009;<&#x2009;0.01, 37.1% vs. 2.9%, respectively). With full informed consent, 20 patients without euploidy from normal fertilization chose 1PN-derived biparental and diploid blastocysts to transfer, resulting in 10 clinical pregnancies and 9 ongoing pregnancies. LARGE-SCALE DATA: N/A. LIMITATIONS REASONS FOR CAUTION: Some rare ploidy abnormalities, such as polyploidy with an equal number of identical sets of chromosomes and ploidy mosaicism cannot be accurately identified. Moreover, the origin of abnormal ploidy was not identified due to the unavailability of DNA from both parents. WIDER IMPLICATIONS OF THE FINDINGS: The PGT-Plus AI model provides a ploidy evaluation method based on the conventional PGT-A data and integrates directly into standard PGT-A workflows. Clinical utility results suggest that the model is a valuable tool for identifying embryos with abnormal ploidy in PGT-A and rescuing normal diploid embryos from abnormally fertilized embryos. These findings demonstrate that PGT-Plus significantly enhances the diagnostic accuracy of PGT. STUDY FUNDING/COMPETING INTERESTS: This study was supported by grants from Major Scientific Program of CITIC Group (No. 2023ZXKYB34100, to Ge.L.), Hunan Provincial Grant for Innovative Province Construction (2019SK4012), Hunan Xiangjiang New District (Changsha High-tech Zone) key core technology research project in 2023, and Science Foundation of Hunan Province (Grant 2023JJ30422). All authors declared no conflicts of interest..

artificial intelligence

Scaling linear-model breeding values to the liability scale: an application to pig binary traits.

In commercial pig production, many important traits are recorded as binary phenotypes. For such traits, threshold models offer an appropriate framework but are computationally intensive. Thus, linear models are widely used to obtain genomic estimated breeding values (GEBV); however, these are on the observed scale (phenotypic). This creates the need for a robust method to approximate GEBV from linear models to the liability scale. A recently proposed approximation showed good concordance for low-prevalence traits (<5%) but has not yet been tested for a wider range of prevalence values and for models with more than one random effect. We aimed to evaluate the performance of this approximation for pig binary traits with prevalences ranging from <5% to >86%, in both animal and maternal animal models. Data were available for five fitness traits (FT1-FT5), with up to 233k animals with phenotypes, of which 204k animals were genotyped with a 25k SNP array. Variance component estimates were obtained using threshold models. Classical animal models were used for FT1-FT3, and maternal animal models for FT4 and FT5. Variance components on the observed scale were then obtained by multiplying estimates from a threshold model by the square of the height of the standard normal density evaluated at the threshold. GEBV were predicted using single-step genomic best linear unbiased prediction under both linear and threshold models. The approximation tested involved scaling the GEBV using the height of the ordinate of the standard normal distribution evaluated at the threshold as a scaling factor. The agreement between GEBV from the scaled linear model and the threshold model on the probability scale was evaluated using Pearson and Spearman correlations, mean squared error (MSE), regression parameters, overlapping coefficient (OVL), distribution overlap, and classification accuracy (CACC). Correlations between linear and threshold GEBV ranged from 0.94 (low-prevalence traits) to 0.99 (high-prevalence traits) for the direct GEBV and were 0.99 for the maternal GEBV. MSE were close to zero. The OVL exceeded 0.83 for all traits. CACC ranged from 95.10% to 98.33% for the direct GEBV and from 92.54% to 97.42% for the maternal GEBV. Regardless of model and trait prevalence, this approximation yielded GEBV that are highly consistent with threshold model GEBV, providing a reliable, practical approach for large-scale pig genetic evaluations for binary traits using linear models.

Animals

Spectral Transforms as a Tool to Optimize Digital Phenotyping in Biological Images.

Modern livestock breeding has mastered genotyping. Genome-wide association studies, genomic selection, and SNP arrays enable genetic merit prediction at lower cost. However, phenotyping remains the bottleneck, as manual measurement is slow, expensive, subjective, and unable to capture spatial or temporal trait organization. Digital phenotyping via artificial intelligence could resolve this, but deep learning requires thousands of labelled examples, impractical when phenotyping cost itself limits datasets to hundreds of individuals. This creates a paradox: AI could accelerate phenotyping but requires large numbers of samples to train the models. Here, we demonstrate that integrating computer vision with machine learning offers sample-efficient digital phenotyping using eggshell colour as a model system. Rather than learning features from scratch (deep learning), we engineer physically motivated features via Wavelet transforms that decompose images into multi-scale spatial components. Wavelet features captured 14.2 percentage points more variance (R2&#x2009;=&#x2009;0.976 vs. 0.834, p&#x2009;<&#x2009;0.001) than standard colorimetry, with 50% better sample efficiency (achieving at n&#x2009;=&#x2009;60 what colorimetry required n&#x2009;=&#x2009;120). Variance decomposition revealed 77% of discriminative capacity derives from spatial patterns (bands, spots, gradients) invisible to scalar averages. Additionally, we identified "cryptic phenotypes" (3.3%) where spatial patterns contradicted average colour, cases where colorimeters failed but Wavelets succeeded. The underlying principle-that spatial decomposition can recover organizational information lost by scalar averaging-may be applicable to other traits with spatial or temporal structure, such as marbling, dermatitis, or pigmentation rhythms, although whether comparable performance gains would be observed remains to be tested empirically. Hence, for breeding programs implementing genomic selection, computer vision-based digital phenotyping captures complex trait variation without massive training datasets, addressing the bottleneck that increasingly limits genetic progress as genotyping becomes trivial.

Wavelet transform

Identification and fine mapping of a locus controlling multi-main-stem trait in Brassica napus.

BACKGROUND: The main stem is a crucial component determining individual plant yield in rapeseed (Brassica napus). However, the genetic and developmental basis underlying the multi-main-stem trait remains largely unclear. RESULTS: In this study, we identified a multi-main-stem mutant, mms1, which exhibited a significantly increased silique number per plant and abnormal shoot apical meristem (SAM) development. Genetic analysis demonstrated that the multi-main-stem trait was controlled by a recessive gene. Using bulked segregant analysis combined with a Brassica napus 50&#xa0;K SNP array and map-based cloning, the locus was mapped to a 340-kb interval on chromosome A09 of the ZS11 reference genome and was designated BnaA09.MMS1. Candidate gene analysis revealed that BnaA09G0254500ZS, which harbors sequence variations in both the promoter and coding regions and shows significantly increased expression in the mutant, was the most likely candidate gene. In addition, phytohormone analysis revealed reduced auxin accumulation in mutant SAMs, together with transcriptomic changes in genes associated with the CLAVATA3 (CLV3)-WUSCHEL (WUS) feedback loop. CONCLUSIONS: These findings provide an important foundation for elucidating the genetic basis of the multi-main-stem trait and offer a valuable genetic resource for rapeseed improvement.

Brassica napus