Search PubMedSearch

SEARCH · Search PubMed

Results for “Admixture model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

696 recordsLinked to original sources

Statistical test to compare the linkage model and the admixture model based on central limit results.

In the Admixture Model, the probability that an individual carries a certain allele at a specific marker depends on the allele frequencies in K ancestral populations and the proportion of the individual's genome originating from these populations. The markers are assumed to be independent. The Linkage Model is a Hidden Markov Model that extends the Admixture Model by incorporating linkage between neighboring loci. We prove consistency and asymptotic normality of maximum likelihood estimators for the ancestry of individuals in the Linkage Model, complementing earlier results by (Pfaff et al., 2004; Pfaffelhuber and Rohde, 2022; Heinzel, 2025) for the Admixture Model. These results are used to prove that a statistical test that allows for model selection between the Admixture Model and the Linkage Model is an asymptotic level-α-test. Finally, we demonstrate the practical relevance of our results by applying the test to real-world data from The 1000 Genomes Project Consortium (2015).

Genetic Linkage

Quo vadis, BGA? A collaborative EDNAP exercise on the challenges and progress in forensic biogeographical ancestry inference.

There is a broad consensus that forensic tests for the prediction of externally visible characteristics (EVC) and analysis of biogeographic ancestry (BGA) of an individual are technically reliable. However, interpretation of the results and population-specific genotype distribution patterns remains challenging. EVC and BGA analyses provide valuable information for population genetics studies and as investigative leads for criminal cases, as well as for historical and contemporary identification tests. However, inaccurate or incorrect predictions, for example, from subjective bias in the interpretations made, have the potential to misdirect police investigations. The legal situation regarding EVC and BGA testing varies by country: ranging from countries where it is explicitly prohibited, to those without specific regulations on biogeographic ancestry prediction, and others that have already enacted laws governing its use. The reluctance to utilize these analyses is not only due to legal restrictions and data protection concerns, but also to initial limited sets of sufficiently comprehensive forensic DNA assays. Forensic BGA marker panels typically contain up to ∼300 SNPs. This relatively small number of genetic markers, along with limited reference population data, complicates the interpretation of results from donors of unknown origin. This paper presents the results of a collaborative EDNAP study, which, for the first time, evaluated the approach to reporting EVC and BGA data between international laboratories. For the study, DNA from nine individuals with self-reported ancestry was collected and analysed using various forensic panels differing in the number and composition of ancestry-informative markers genotyped, comprising: the Precision ID mtDNA Whole Genome Panel, the VISAGE Basic Tool and the VISAGE Enhanced Tool for Appearance and Ancestry Prediction, and the Ion AmpliSeq™ PhenoTrivium Panel. To ensure full data protection, all SNP genotypes and uniparental marker haplotypes obtained were not shared with third parties. Instead, the genetic data were analysed using a range of commonly used population analysis software packages. These analysis outcomes were then distributed to twelve European forensic laboratories (both academic and law enforcement institutions), who were asked to prepare reports based on their interpretation of the phenotypes and ancestry they inferred from the analysis data. A questionnaire sent alongside the genetic information, aimed to evaluate which difficulties were encountered by the participants in processing the BGA analysis data they were given.

Humans

Genomic history of the Caucasus: A systematic review and meta-analysis of ancient DNA studies.

The Caucasus region represents a unique natural laboratory for paleogenetic research due to its complex topography, long-standing role as a migratory corridor and glacial refugium, and exceptional preservation conditions for ancient DNA. This review synthesizes recent genome-wide studies to reconstruct the demographic history shaping the distinctive genetic landscape of modern Caucasus populations. The analysis reveals a deep pattern of continuity, isolation, and periodic admixture. Early genetic differentiation emerged in the Neolithic and Chalcolithic, forming distinct steppe and mountain population clusters. The Bronze Age was a pivotal period marked by large-scale gene flow from the Eurasian Steppe, particularly linked to the Yamnaya expansion, and interactions with Iranian and Anatolian-related groups. Despite these influences, many populations demonstrate remarkable genetic continuity from the Bronze Age to the present day. Significant knowledge gaps persist, particularly for the Paleolithic, Mesolithic, and Neolithic of the North Caucasus, as well as for the Late Medieval and Early Modern periods across the entire region. Addressing these gaps through targeted archaeogenomic studies is crucial for understanding the fine-scale processes that formed the hierarchical structure and high linguistic diversity of Caucasus populations, offering a powerful model for studying human adaptation, interaction, and language-genetics dynamics in a mountainous environment.

Humans

Genome-wide SNP data support species boundaries in sympatric Polylepis Ruiz & Pav. (Rosaceae) species from Bolivia and Ecuador.

Species delimitation in the South American genus Polylepis is notoriously challenging due to high morphological similarity and phenotypic plasticity, likely driven by hybridization and gene flow. Previous phylogenetic studies suggested that genetic structure aligns more strongly with geography than with taxonomy, questioning existing species concepts and hampering conservation efforts. We used double-digest RAD sequencing (ddRADseq) to generate genome-wide SNP data for 11 Polylepis species sampled across multiple localities in Bolivia and Ecuador. Population genetic analyses, phylogenetic inference, and network approaches were combined to assess whether genetic structure aligns more closely with taxonomy or geography. Morphologically defined species formed largely cohesive genetic lineages across regions, with species identity explaining substantially more genetic variation than locality. While localized admixture and reticulation were detected among closely related taxa, widespread species showed strong genetic cohesion and clear separation from congeners. Our results indicate that the sampled Polylepis species from Bolivia and Ecuador maintain distinct genetic identities despite localized signals consistent with gene flow. This genome-wide support for current taxonomy highlights Polylepis as a valuable model for studying speciation under gene flow and indicates that multiple geographic sampling will be essential in reconstructing a robust phylogeny of the genus, with important implications for conservation planning in Andean montane forests.

Bolivia

Genome-wide scans reveal candidate genes associated with wing morph differentiation in Tetrix japonica.

Wing dimorphism is an important dispersal-related trait in insects, but its genomic basis remains poorly understood in pygmy grasshoppers. Here, we integrated genome-wide single-nucleotide polymorphism (SNP) analyses, population structure inference, selection scans, and functional annotation to investigate genomic differentiation between long- and short-winged Tetrix japonica. Principal component analysis (PCA), ADMIXTURE, and phylogenetic analyses revealed weak genome-wide separation between morphs, indicating differentiation on a largely shared genetic background. Genome-wide scans based on the fixation index (FST), nucleotide diversity ratios, and Tajima's D, using 50-kb non-overlapping windows and empirical top-5% outlier thresholds, identified multiple candidate regions across seven chromosomes. The broader long- and short-winged candidate sets spanned 9.35 Mb and 9.37 Mb and directly overlapped 82 and 77 genes, respectively. Candidate genes were associated with signaling/hormone regulation, membrane transport, metabolism, cytoskeletal organization, extracellular matrix structure, and development. Short-winged candidate genes were significantly enriched for ABC-type transporter activity and ATP hydrolysis activity. Because all individuals originated from a single laboratory-maintained population with weak genome-wide structure, these regions should be regarded as candidate loci from a screening-stage analysis that require validation in independent populations and by functional assays, rather than as confirmed targets of selection.

Animals

Modelling the effects of biological intervention in a dynamical gene network.

Cellular response to environmental and internal signals can be modeled by dynamical gene regulatory networks (GRN). In the literature, three main classes of gene network models can be distinguished: (1) non-quantitative (or data-based) models which do not describe the probability distribution of gene expressions; (2) quantitative models which fully describe the probability distribution of all genes co-expression; and (3) mechanistic models which allow for a causal interpretation of gene interactions. We propose two rigorous frameworks to model gene alteration in a dynamical GRN, depending on whether the network model is quantitative or mechanistic. We explain how these models can be used for design of experiment, or, if additional alteration data are available, for validation purposes or to improve the parameter estimation of the original model. We apply these methods to the Gaussian graphical model, which is quantitative but non-mechanistic, and to mechanistic models of Bayesian networks and penalized linear regression.

Gene Regulatory Networks

The landscape of pruning for large language models: A systematic review and unified taxonomy.

Confronting the inherent tension between the exceptional capabilities and the immense computational costs of Large Language Models (LLMs), pruning has become a crucial technique for achieving efficient deployment. However, a systematic analytical framework dedicated specifically to LLM pruning remains absent. In this paper, we aim to bridge this gap. We first elucidate the theoretical foundations that underpin the effectiveness of pruning, namely overparameterization and redundancy, and then propose a multidimensional taxonomy that organizes existing approaches along the axes of granularity, timing, and criteria. Building upon this unified perspective, we further analyze performance recovery mechanisms and the broader evaluation ecosystem, while also exploring forward-looking challenges such as interpretability, automation, and hardware-algorithm co-design. Through this comprehensive synthesis, we seek to provide an integrated and coherent analytical lens for advancing both research and practice in LLM pruning.

Large Language Models

An integrated multiscale air quality modelling framework for industrial park pollution: Linking local emissions to regional transport.

Capturing the spatiotemporal distribution of pollutants in industrial parks remains challenging for regional air quality models because of their coarse resolution (3 km), resulting in uncertainties in local emission quantification. To address this, we developed the Integrated Multiscale Air Quality Modelling System for Industry (IAQMS-Industry), coupling the regional Nested Air Quality Prediction Modelling System (NAQPMS) with a city-scale chemical transport model. This framework integrates point-source locations and Gaussian plume dispersion to simulate particulate matter with a diameter smaller than 2.5 micrometres (PM2.5) at 100 m resolution. Applied to the Beijing Yi Zhuang and Tangshan industrial parks and evaluated against observations. The coupled model achieved a normalized mean bias (NMB) ranging from 3.1 % to 6.2 %, improving upon NAQPMS (-16.9 % to -7.7 %). Spatial analysis revealed that coarse regional grids underestimated the PM2.5​ concentrations at industrial sites by smoothing gradients, whereas IAQMS-Industry successfully resolved spatial patterns. Industrial point emissions accounted for 22.9 %-26.4 % of PM2.5 in the coupled model, which was significantly greater than the regional model estimates of 1.6 %-13.7 %. These findings indicate that regional models overestimate pollutant dispersion processes in industrial parks while underestimating local industrial impacts. By explicitly resolving point-source dynamics and linking them to regional transport, IAQMS-Industry provides a robust tool for designing targeted emission controls in industrial cities and balancing local air quality improvements with minimized regional pollution outflow. This study underscores the necessity of multiscale modelling for accurate source apportionment and informed environmental governance in industrial zones.

Air Pollution

Penalized Cumulative Probability Model for a Continuous Outcome Subject to Detection Limits.

Mixed-type outcome data occur when the outcome variable's distribution is a mixture of both continuous and discrete ordinal variables. Such mixed-type outcomes are common in biomedical, psychological, and the health sciences, particularly for variables having either a detection or quantitation limit. When interest lies in identifying a combination of genomic features associated with a mixed-type outcome, any method used would require a variable selection strategy for high-dimensional data. Unfortunately, few variable selection methods exist for modeling a mixed-type outcome when the covariate space is high dimensional. This study develops a high-dimensional penalized cumulative probability model (CPM), to allow for the identification of genomic features associated with mixed-type outcome of interest. We demonstrated how such model may be estimated using the iterative penalization procedure-the generalized monotone incremental forward stagewise (GMIFS) algorithm. The Model-X knockoffs procedure was combined with the estimation algorithm to control the false discovery rates (FDR) when performing variable selection. Through extensive simulation studies, our penalized CPM was shown to outperform alternative methods in terms of controlled variable selection performance by achieving high statistical power with the FDR being controlled at the target level. We demonstrate the utility of our method by applying it to predict estimated glomeruli filtration rate (eGFR) in kidney transplant recipients at 24 months post-transplant using baseline gene expression data as predictors. Our CPM model identified five genes associated with this mixed-type outcome which have important links to renal disease, which may provide prognostic guidance for kidney transplantation recipients.

Models, Statistical

Construction of precision clinical-proteomics risk model based on machine learning for predicting heart failure in type II diabetes mellitus.

BACKGROUND AND AIMS: Heart failure (HF) is a severe complication in type 2 diabetes mellitus (T2DM), but current risk stratification scores have limited predictive accuracy. We aimed to develop novel prediction tools integrating clinical variables with proteomics to improve risk stratification of hospitalization for HF in T2DM. METHODS AND RESULTS: In this study, we included 2111 UK Biobank participants with T2DM but no prior HF, and profiled 2920 proteins to predict 10-year incident HF hospitalization. Participants were randomly divided into training (70%), tuning (10%), and validation (20%) sets.Three prediction models were developed: a Clinical model based on demographic characteristics, comorbidities, medication use, and laboratory indices; a Protein model based on 40 proteins selected by the Light Gradient Boosting Machine (LGBM); and the Clinical OMics and Protein ASSessment for Heart Failure (COMPASS-HF) model, which integrated both clinical variables and the LGBM-selected proteins. Models were evaluated for area under the curve (AUC), sensitivity, and specificity. During follow-up, 168 participants (7.96%) developed incident HF. The COMPASS-HF model showed better discrimination than the Clinical model, with an AUC of 0.897 (95% CI: 0.850-0.945) versus 0.790 (95% CI: 0.723-0.856). It also demonstrated higher sensitivity (0.882; 95% CI: 0.725-0.967) and consistent performance in subgroups. COMPASS-HF effectively stratified risk of hospitalization for HF, with cumulative incidence rates of 31.9% in the high-risk group and 1.2% in the low-risk group. CONCLUSIONS: By combining clinical and proteomic variables, we developed a high-performance HF prediction model for T2DM, enabling precise risk stratification and informing early intervention strategies.

Humans

Beyond predictive performance: A systematic review and critical methodological appraisal of AI/ML and conventional modelling strategies in breast, colorectal, and pancreatic Cancer.

BACKGROUND: Predictive modelling for cancer risk, treatment-related complications, and survival is central to precision oncology. Conventional logistic regression (LR) and Cox proportional hazards (CoxPH) regression remain widely used but are limited when modelling nonlinear interactions, high-dimensional imaging features, and multimodal clinical-metabolic predictors. Artificial intelligence (AI) and machine learning (ML) methods offer expanded capability through automated feature extraction, ensemble learning, and flexible survival modelling, but the evidence on when AI/ML adds value over conventional models across cancer sites and predictive tasks remains fragmented. OBJECTIVE: To systematically evaluate the methodological performance, validation strategies, and translational limitations of AI/ML models compared with conventional statistical models in published predictive-modelling studies for breast, colorectal, or pancreatic cancer. METHODS: PubMed, Scopus, and Web of Science were searched for studies published between January 2019 and March 2025. Two reviewers independently conducted title-and-abstract screening, full-text eligibility assessment, and PROBAST risk-of-bias assessment. Sixty-five studies (n = 907,567 participants) were narratively synthesised by cancer site, predictive task, model family, comparator, validation strategy, predictor modality, and calibration or explainability reporting. RESULTS: The 65 studies comprised breast cancer (n = 35), colorectal cancer (n = 21), and pancreatic cancer (n = 9). AI/ML superiority over LR and CoxPH was task- and data-dependent. CNN- and U-Net-based models predominated in imaging and body-composition tasks, tree-based ensembles consistently outperformed LR for tabular perioperative complication prediction, and CoxPH remained competitive, and in the largest pancreatic risk study, superior to XGBoost (C-index 0.802 vs 0.723) in well-structured datasets. PROBAST analysis-domain risk was moderate in 54 of 65 studies (83%), driven by limited external validation, sparse calibration reporting (11/65), and few decision-curve analyses (7/65). CONCLUSION: AI/ML adds the most methodological value in imaging-derived feature extraction and nonlinear perioperative prediction, while conventional regression remains preferable in large, structured datasets with linear predictors. Clinical translation requires standardised body-composition definitions, external validation, calibration assessment, decision-curve analysis, and explainability, in line with TRIPOD+AI and CLAIM standards.

Humans

Revealing the Shared Genetic Architecture of Metabolic Dysfunction-Associated Steatotic Liver Disease-Related Traits Through Genomic Structural Equation Modeling.

Although individual traits related to metabolic dysfunction-associated steatotic liver disease (MASLD) have been investigated through large-scale genome-wide association studies (GWASs), the shared genetic susceptibility across these traits remains unclear. We therefore conducted a multivariate GWAS of key MASLD-related traits to elucidate their common genetic architecture. We applied genomic structural equation modeling to model a latent genetic factor (MASLD-F) underlying genetically correlated MASLD-related traits, leveraging their GWAS-derived genetic correlations. We then performed functional annotations, including fine-mapping, transcriptome-wide association study, and cell- and tissue-type-specific enrichment analyses, and conducted Mendelian randomization analyses to identify modifiable risk factors. Our multivariate MASLD-F GWAS identified 50 independent variants across 48 genomic loci. Transcriptomic imputation identified several MASLD-F-associated genes, including ARNTL, NPC1, BTBD10, VDAC2, TSKU, SFMBT1, and ABHD17C. We observed significant enrichment of MASLD-F-related genetic signals predominantly in brain tissues, pancreatic islets, and the adrenal gland. Additionally, six modifiable risk factors and four modifiable protective factors for MASLD-F were identified. These findings reveal a complex shared genetic architecture underlying MASLD components, thereby expanding our understanding of disease pathogenesis and providing novel insights for precision medicine and public health interventions.

Humans

Evaluation of a cornea-specialized large language model for diagnostic and management accuracy in complex corneal cases.

PURPOSE: To evaluate whether a cornea-specialized large language model (LLM) enhanced with retrieval-augmented generation (RAG) improves clinicians' diagnostic and management accuracy in complex corneal cases compared to a general-purpose GPT-4o model and unaided clinician performance. METHODS: This prospective, randomized, masked evaluation study involved three cornea trainees who each independently reviewed 39 real-world corneal cases under three experimental conditions: unaided, GPT-4o-assisted, and assisted by a cornea-specialized GPT-4o model. The cornea-specialized model was constructed by embedding over 200 publicly available Wikipedia articles into GPT-4o's RAG framework. Participants provided open-ended diagnoses and selected the next-step management options (multiple choice). They were allowed up to three GPT-4o queries per case, and the AI-assisted arms were randomized to minimize bias. Accuracy for both tasks was compared against expert reference standards using McNemar's test. RESULTS: Diagnostic accuracy was 48.7%, 20.5%, and 38.5% unaided, improving to 69.2%, 46.2%, and 59.0% with general GPT-4o (p<0.04). The cornea-specialized GPT-4o further improved accuracy to 71.8%, 48.7%, and 74.4%, with improvements over unaided performance for all clinicians (p<0.01). For next-step decisions, unaided accuracy was 76.9%, 87.2%, and 59.0%. With the specialized model, Ophthalmologist 3 improved to 71.8% (p<0.05), Ophthalmologist 1 remained high at 82.1%, and Ophthalmologist 2 declined to 64.1% (p<0.05). CONCLUSIONS: A cornea-specialized LLM enhanced with RAG improved diagnostic accuracy in complex corneal cases, particularly among clinicians with lower baseline performance. Effects on management accuracy were inconsistent. Future studies should explore the use of open-ended management tasks and examine whether smaller, curated retrieval corpora yield better model performance.

Humans

Rational design of high-productivity perfusion processes for CHO Cells: From growth inhibitory strategies to model-driven optimization.

While perfusion culture for Chinese hamster ovary (CHO) cells offers advantages such as continuous operation and flexibility, it suffers from product loss through cell bleeding and difficulties in reaching high productivity due to sustained rapid cell growth. Growth inhibitory strategies are widely used to enhance productivity in fed&#x2011;batch processes; however, their practical implementation and comparative effectiveness in perfusion processes remain insufficiently explored. Meanwhile, process development often relies on costly trial&#x2011;and&#x2011;error approaches. Here, we systematically compared three growth inhibitory strategies in perfusion culture-low cell&#x2011;specific perfusion rate (CSPR), sodium butyrate, and mild hypothermia-with respect to cell growth, metabolism, productivity, and product quality. Genome&#x2011;scale metabolic flux sampling analysis revealed that low&#x2011;CSPR and sodium butyrate induce a convergent up&#x2011;regulation of energy metabolism, correlating with greater gains in specific productivity (qp). Building on this insight, we developed a growth&#x2011;kinetic model for the combined low&#x2011;CSPR + butyrate strategy, incorporating parameter uncertainty. This model&#x2011;guided framework enabled the rational design of two distinct high&#x2011;productivity perfusion processes: a sustained mode that achieved robust long&#x2011;term stability alongside substantial productivity gains, and a high&#x2011;intensity mode that pushed qp and daily volumetric titer to their maxima, with increases of up to 108.94% and 190.36%, respectively, in a model CHO cell line with a moderate baseline productivity. Our study provides a proof&#x2011;of&#x2011;concept framework for perfusion intensification, from strategy selection to rational process design.

Animals

Risk prediction models for blood transfusion in patients undergoing total hip and knee arthroplasty: a systematic review and meta-analysis.

OBJECTIVE: To systematically review and evaluate published risk prediction models for perioperative blood transfusion in patients undergoing total hip or knee arthroplasty (THA/TKA). METHODS: We systematically searched PubMed, Web of Science, the Cochrane Library, and Embase from inception to May 31, 2025. Two researchers independently screened the literature, extracted data, and assessed the risk of bias and applicability using the Prediction model Risk Of Bias Assessment Tool (PROBAST). The area under the receiver operating characteristic curve (AUC) values were pooled via a meta-analysis using Stata 18.0. RESULTS: d Fourteen studies containing 36 prediction models were included. The incidence of blood transfusion among THA/TKA patients ranged from 3.2% to 30.8%. Preoperative hemoglobin (Hb) level, tranexamic acid (TXA) use, operative duration, intraoperative blood loss, and age were the most frequently incorporated predictors. Model sensitivity ranged from 58% to 94.5%, and specificity ranged from 71.3% to 94%. Meta-analysis showed that the pooled AUC value of the 13 validated models was 0.87 (95% CI: 0.85-0.90), suggesting good discriminatory performance. All models were rated as having a high risk of bias. The applicability of four studies was rated as unclear. CONCLUSION: Although the included studies demonstrated promising discriminative ability of prediction models for blood transfusion in THA/TKA, all were assessed as having a high risk of bias using the PROBAST tool. Therefore, future research should prioritize the development of models with larger sample sizes, rigorous study designs, and multicenter external validation.

Humans

Development and Validation of a Predictive Model for Identification of Cognitive Impairment Risk in Older Adults with Subjective Cognitive Decline&#xff1a;A Longitudinal Study.

BACKGROUND: Subjective cognitive decline (SCD) is a transitional state between objective cognitive impairment and cognitively intact mental status, providing a critical window for implementing preventive interventions to delay objective cognitive decline. AIMS: We aimed to develop a predictive model for SCD progression in older adults with mild cognitive impairment (MCI). This model will facilitate the identification of risk factors and establishment of targeted interventions for community-based SCD management. METHODS: Data from the China Health and Retirement Longitudinal Study (CHARLS) was utilized in this study, extracting 18 indicators. Potential predictors selected through univariate Cox regression and LASSO regression analyses were sequentially incorporated into a multivariable Cox regression model. A nomogram was constructed to establish a predictive model. Model validation encompassed Area Under Curve (AUC) metrics for discriminative capacity, complemented by quantitative assessments using calibration curve analysis for precision verification and decision curve analysis (DCA) for clinical utility evaluation. RESULTS: A total of 1099 older adults with SCD were included in the final analysis, of whom 114 (10.3%) developed MCI. Multivariable Cox regression identified residence, marital status, educational level, social participation, gait speed, and baseline cognitive function. The model demonstrated time-dependent AUC values of 0.885, 0.830, 0.839, and 0.836 in the training set when evaluating discriminative capacity at 2-, 4-, 7-, and 9-year, respectively. The predictive model showed excellent predictive ability according to AUC, calibration curve, and DCA. CONCLUSIONS: A predictive model was created to estimate the risk of developing MCI in older individuals with SCD, offering clinician-actionable intervention benchmarks for preventive care.

Humans