Search PubMedSearch

SEARCH · Search PubMed

Results for “Machine learning model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

AI-Supported, Integrative Prediction of Postoperative Delirium: Protocol for the CONFUSED Study.

BACKGROUND: Postoperative delirium (POD) is a frequent and serious complication in older surgical patients, characterized by acute cognitive dysfunction and fluctuating levels of consciousness. POD is associated with prolonged hospitalization, long-term cognitive decline, reduced quality of life, and increased mortality. Despite its clinical relevance, the underlying pathophysiological mechanisms remain poorly understood, and reliable biomarkers for early prediction and prevention are lacking. OBJECTIVE: The CONFUSED study aims to identify molecular and clinical predictors of POD by integrating clinical data with proteomic, transcriptomic, and epigenetic analyses. The primary objective is to develop predictive models for POD using multimodal data. Secondary objectives include the identification of delirium-associated genes, proteins, and epigenetic signatures, as well as the exploration of patient subgroups at increased risk for POD. METHODS: CONFUSED is a prospective observational cohort study conducted at a German university hospital. Adult patients undergoing major surgery under general anesthesia will be enrolled until 100 cases of POD have been observed, which is expected to require a total sample size of approximately 200 to 300 patients. Blood samples are collected at 4 predefined time points: before premedication, immediately after surgery, and on postoperative days 2 and 5. Samples undergo comprehensive proteomic profiling, transcriptomic analysis using RNA microarrays, DNA methylation analysis, and genotyping of selected polymorphisms. Clinical data, including demographics, comorbidities, perioperative variables, medications, and delirium assessments using the Confusion Assessment Method (CAM) and CAM for the intensive care unit, are systematically recorded. Statistical analyses include univariate and multivariate methods, as well as machine learning approaches such as random forests and support vector machines, to identify relevant biomarkers and develop predictive models. The study protocol follows STROBE (Strengthening the Reporting of Observational Studies in Epidemiology) and TRIPOD (Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis) guidelines and was approved by the responsible ethics committees. RESULTS: The study was registered in the German Clinical Trials Register (DRKS00033854) on March 18, 2024. Recruitment started in January 2024 and is ongoing at the time of manuscript submission. As of now, 135 patients have been enrolled. Sample collection and laboratory analyses are ongoing. Data analysis began in January 2026, with first results anticipated in July 2026. Final data lock is anticipated after the completion of recruitment. CONCLUSIONS: By integrating multimodal molecular data with clinical parameters and applying advanced machine learning techniques, the CONFUSED study aims to improve the prediction and understanding of POD. The results are expected to support the development of personalized preventive strategies and contribute to improved perioperative care for patients at risk of POD.

Humans

Exploring the mechanism of aroma production in fermented cherry juice by L. brevis LD1.0600 using flavomics and whole genome analysis.

This study focused on L.brevis LD1.0600 with excellent fermentation traits: it analyzed genome-wide key regulatory genes for micro-metabolites, combined with fermented cherry juice flavor metabolomics data, and used machine learning to explore correlations between gene regulation, metabolite production, and flavor formation. The SVM model screened and verified fermented cherry juice VOCs; through OAV and flavor wheel analysis, LD1.0600 emerged as the top-performing strain, with a sweet, fruity dominant aroma. Key aroma-active components (OAV > 100) included 2-methoxy-4-vinylphenol, benzaldehyde, 2-methyl-butanoic acid and hexanoic acid, and 2-methoxy-4-vinylphenol and hexanoic acid elevated by LD1.0600-regulated genes (Chrom1-001884, Chrom1-000925, fabF and Chrom1-000199). At the same time, through research, a "strain screening-SVM screening of DVCs-OAV screening of key aroma components-whole genome sequencing of flavor regulatory genes" system was established. This system can not only be applied to the screen fermentation strains, but also can be extended to the application of other fermentation products.

Fermentation

A robust transfer learning approach for high-dimensional linear regression to support integration of multi-source gene expression data.

Transfer learning aims to integrate useful information from multi-source datasets to improve the learning performance of target data. This can be effectively applied in genomics when we learn the gene associations in a target tissue, and data from other tissues can be integrated. However, heavy-tail distribution and outliers are common in genomics data, which poses challenges to the effectiveness of current transfer learning approaches. In this paper, we study the transfer learning problem under high-dimensional linear models with t-distributed error (Trans-PtLR), which aims to improve the estimation and prediction of target data by borrowing information from useful source data and offering robustness to accommodate complex data with heavy tails and outliers. In the oracle case with known transferable source datasets, a transfer learning algorithm based on penalized maximum likelihood and expectation-maximization algorithm is established. To avoid including non-informative sources, we propose to select the transferable sources based on cross-validation. Extensive simulation experiments as well as an application demonstrate that Trans-PtLR demonstrates robustness and better performance of estimation and prediction when heavy-tail and outliers exist compared to transfer learning for linear regression model with normal error distribution. Data integration, Variable selection, T distribution, Expectation maximization algorithm, Genotype-Tissue Expression, Cross validation.

Linear Models

Unveiling tumor heterogeneity by single cell RNA-sequencing: From basic considerations to clinical applications.

Tumor heterogeneity-encompassing diverse cellular phenotypes, genomic alterations, and microenvironmental contexts-is a principal barrier to effective cancer therapy. Single-cell RNA sequencing (scRNA-seq) has transformed our ability to resolve this complexity by capturing transcriptomes at single-cell resolution. Here, we review the technical foundations required for high-quality scRNA-seq studies. We then trace the evolution of scRNA-seq platforms from manual micromanipulation to high-throughput systems, and describe the computational pipelines that enable reliable data interpretation. The application of scRNA-seq is exemplarily shown in the context of lung cancer, where single-cell profiling has revealed (i) the clonal and sub-clonal architecture of tumors, (ii) extensive remodeling of the immune microenvironment, iii) key mechanisms underlying resistance to targeted agents and immune-checkpoint blockade, and (iv) the dynamics of neo-antigen-specific T-cell responses. Integrating machine-learning techniques-such as deep-learning classifiers and graph-based models-with single-cell transcriptomic data has markedly sped up biomarker discovery, produced more accurate risk-stratification scores, and enabled the generation of patient-specific therapeutic predictions. We surveyed the major trial registry ClinicalTrials.gov and identified ∼380 ongoing or completed studies that explicitly incorporate scRNA-seq as a correlative or pharmacodynamic endpoint. Overall, the analysis shows that scRNA-seq becomes an increasingly important component of modern trials, providing high-resolution cellular and molecular readouts that complement conventional imaging and bulk-omics endpoints. While key challenges remain, ranging from costs, scalability and need for rigorous validation before routine clinical deployment, ongoing technological advances continue to expand the potential of scRNA-seq as a cornerstone of precision medicine.

Humans

Profiler: an open web platform for multi-omics analysis.

MOTIVATION: High-throughput multi-omics technologies produce increasingly large and heterogeneous datasets that are difficult to analyze without advanced computational expertise. Existing bioinformatics tools are often fragmented or limited to specific omics types, hindering reproducibility and accessibility. There is a critical need for an integrated, user-friendly, and scalable platform capable of supporting multi-omics analyses across different data modalities. RESULTS: We present Profiler, an open-source, modular platform that unifies data import, quality control, preprocessing, statistical testing, machine and deep learning, biomarker discovery, pathway and drug-target enrichment, and survival modeling within a single reproducible environment. Built in Python with Streamlit, Profiler is available as both a web-based platform deployed on high-performance computing and a desktop version for local execution, enabling flexible usage across computational infrastructures. Profiler supports diverse omics modalities, including proteomics, transcriptomics, lipidomics, and electroencephalogram data. Through applications to glioblastoma proteomic, pancancer, and multi-omics datasets, Profiler reproduced known molecular subtypes, revealed potential therapeutic targets, and generated fully traceable analysis reports within minutes. By integrating advanced analytics behind an intuitive interface, Profiler democratizes multi-omics analysis and provides a robust, scalable foundation for systems biology and precision medicine research. AVAILABILITY AND IMPLEMENTATION: Profiler is open-source and freely available via its web platform (https://prism-profiler.univ-lille.fr) and GitHub (web version: https://github.com/yanisZirem/Profiler_v1_requests_datatests, desktop version: https://github.com/yanisZirem/prism-profiler), and archived on Zenodo (DOI: https://doi.org/10.5281/zenodo.17478158).

Software

Machine learning-ready genomic biomarkers: ATF3 polymorphisms predict postoperative analgesic demand through AI-compatible phenotyping.

PURPOSE: To determine whether ATF3 polymorphisms can serve as genetic biomarkers for machine learning-based precision analgesia by establishing a genotype-phenotype association suitable for predictive modeling of postoperative opioid requirements. METHODS: In a prospective cohort of 167 adults undergoing abdominal surgery, ATF3 SNPs rs3122721 and rs3125293 were genotyped. A structured dataset architecture was developed to represent genetic profiles as input features for supervised learning models, enabling translational analysis of genotype‑dependent opioid consumption over 72 h. RESULTS: Patients with homozygous genotypes of the ATF3 SNPs had significantly higher opioid requirements than non‑carriers, despite reporting similar subjective pain scores. This consistent genotype‑dependent pattern provided a clinically relevant phenotype suitable for integration into predictive algorithms. CONCLUSION: ATF3 genotyping offers a promising biomarker for computationally informed precision analgesia. By linking genomic variability to clinically meaningful outcomes within a structured clinical and genomic framework, this approach supports the future development of risk-stratified clinical decision-support systems to optimize postoperative pain management.Trial registration ChiCTR1900021991, registered 30 April 2019. SUPPLEMENTARY INFORMATION: The online version contains supplementary material available at https://doi.org/10.1007/s13755-026-00480-9.

ATF3

APNet, an explainable sparse deep learning model to discover differentially active drivers of severe COVID-19.

MOTIVATION: Computational analyses of bulk and single-cell omics provide translational insights into complex diseases, such as COVID-19, by revealing molecules, cellular phenotypes, and signalling patterns that contribute to unfavourable clinical outcomes. Current in silico approaches dovetail differential abundance, biostatistics, and machine learning, but often overlook nonlinear proteomic dynamics, like post-translational modifications, and provide limited biological interpretability beyond feature ranking. RESULTS: We introduce APNet, a novel computational pipeline that combines differential activity analysis based on SJARACNe co-expression networks with PASNet, a biologically informed sparse deep learning model, to perform explainable predictions for COVID-19 severity. The APNet driver-pathway network ingests SJARACNe co-regulation and classification weights to aid result interpretation and hypothesis generation. APNet outperforms alternative models in patient classification across three COVID-19 proteomic datasets, identifying predictive drivers and pathways, including some confirmed in single-cell omics and highlighting under-explored biomarker circuitries in COVID-19. AVAILABILITY AND IMPLEMENTATION: APNet's R, Python scripts, and Cytoscape methodologies are available at https://github.com/BiodataAnalysisGroup/APNet.

COVID-19

Properties Governing Native State Entanglements and Relationships to Protein Function.

Non-covalent lasso entanglements are structural motifs found in a majority of globular proteins, and their misfolding has been linked to a range of biological consequences. Here, we characterize these motifs' structural and physicochemical properties, sequence biases, functional site correlations, and universal features across E. coli, S. cerevisiae, and H. sapiens. We find that the crossing residues, which pierce the plane of the entanglement loop, are 11-times more likely to be a β-strand than an α-helix or random coil, and that around this position the protein sequence is 2.5-times more likely to be composed of a stretch of all hydrophobic residues (most often Val, Ile, or Phe) compared to other sequence motifs. Functionally, crossing residues are enriched at enzyme active sites in S. cerevisiae and small molecule binding residues across all species to degrees greater than expected by random chance. Metal binding residues are enriched in these entanglements in H. sapiens. Increasing statistical power by pooling together these species data, we find RNA-binding residues are enriched in these entanglement components. On the other hand, there is a spatial depletion of crossing residues at sites involved in protein binding. Using machine learning, we identified eight robust features predictive of these entanglements, achieving AUROC scores of 0.8 across species. These results are significant because they suggest a direct role for components of native entanglements in particular protein functions, as well as identifying strong secondary structure and sequence preferences in native entanglements.

Humans

Esketamine multi-omic biomarker evaluation in major depressive disorder (EMBER-MDD): concept, objectives and methodologies of a non-clinical investigator-initiated study.

Treatment resistance (TR) in major depressive disorder (MDD) affects a substantial minority of patients and is hard to recognize early, delaying intensified care. The Esketamine multi-omic biomarker evaluation in MDD (EMBER-MDD) is a non-interventional, investigator-initiated, in-vitro study within the EU Psych-STRATA programme, analyzing biospecimens collected in the randomized INTENSIFY study and the mirror OBS-TR cohort after participants complete treatment. EMBER-MDD aims to discover individual-omic and integrated multi-omic (hypothesis-free) biomarkers and signatures associated with TR risk, and molecular correlates of clinical response to esketamine nasal spray versus treatment as usual (TAU). Biomaterials will derive from approximately 420 adults with MDD (estimated n = 210 esketamine; n = 210 TAU) and include whole blood, RNA-stabilized whole blood, plasma and serum, sampled at baseline and, when feasible, during and after treatment (up to ~ 5,040 aliquots stored at - 80 °C). Genomics will use baseline DNA genotyping on Illumina Infinium GSA v3.0+MD arrays; epigenomics will profile genome-wide DNA methylation across time points using MethylationEPIC v2.0; transcriptomics will employ mRNA-seq (NovaSeq X/ X Plus); and proteomics/ metabolomics will be generated using high-throughput Olink and/ or Biocrates platforms. Each layer will undergo state-of-the-art preprocessing and analyses (e.g., GWAS/ PRS, EWAS, differential expression, WGCNA, pathway and network analyses), followed by integrative strategies including QTL mapping (meQTL/ eQTL/ pQTL/ mQTL) and intermediate-fusion machine learning with nested cross-validation, explainable AI (SHAP/ LIME) and treatment-effect modelling. All outputs are research-only and will not support individual efficacy, tolerability, or clinical decision-making. The study will deliver robust biosignatures and mechanistic hypotheses to guide future validation and inform stratified, molecularly guided intervention strategies in subsequent prospective trials. Trial registration number: 2023-506617-21-00 and 2025-178-f-S.

Humans

Evaluating the impact of modeling choices on the performance of integrated genetic and clinical models.

PURPOSE: The value of genetic information for improving the performance of clinical risk prediction models has yielded variable conclusions. Many methodological decisions have the potential to contribute to differential results. We performed multiple modeling experiments integrating clinical and demographic data from electronic health records with genetic data to understand which decisions may affect performance. METHODS: Clinical data in the form of structured diagnostic codes, medications, procedural codes, and demographics were extracted from 2 large independent health systems, and polygenic risk scores (PRS) were generated across all patients of European ancestry with genetic data in the corresponding biobanks. Crohn's disease was studied based on its substantial genetic component, established electronic health records-based definition, and sufficient prevalence for training and testing. We investigated the impact of choices regarding the PRS integration method, training sample, model complexity, and performance metrics. RESULTS: Overall, our results showed that including PRS resulted in higher performance, but this gain was only robust in situations with limited clinical information. We found consistent performance increases from more compute-intensive models, such as random forest, but the impact of other decisions varied by site. CONCLUSION: This work highlights the importance of considering methodological decision points in interpreting the impact of PRS on prediction performance in clinical models.

Humans

Construction of a molecular diagnostic system for neurogenic rosacea by combining transcriptome sequencing and machine learning.

Patients with neurogenic rosacea (NR) frequently demonstrate pronounced neurological manifestations, often unresponsive to conventional therapeutic approaches. A molecular-level understanding and diagnosis of this patient cohort could significantly guide clinical interventions. In this study, we amalgamated our sequencing data (n = 46) with a publicly accessible database (n = 38) to perform an unsupervised cluster analysis of the integrated dataset. The eighty-four rosacea patients were partitioned into two distinct clusters. Neurovascular biomarkers were found to be elevated in cluster 1 compared to cluster 2. Pathways in cluster 1 were predominantly involved in neurotransmitter synthesis, transmission, and functionality, whereas cluster 2 pathways were centered on inflammation-related processes. Differential gene expression analysis and WGCNA were employed to delineate the characteristic gene sets of the two clusters. Subsequently, a diagnostic model was constructed from the identified gene sets using linear regression methodologies. The model's C index, comprising genes PNPLA3, CUX2, PLIN2, and HMGCR, achieved a remarkable value of 0.9683, with an area under the curve (AUC) for the training cohort's nomogram of 0.9376. Clinical characteristics from our dataset (n = 46) were assessed by three seasoned dermatologists, forming the NR validation cohort (NR, n = 18; non-neurogenic rosacea, n = 28). Upon application of our model to NR diagnosis, the model's AUC value reached 0.9023. Finally, potential therapeutic candidates for both patient groups were predicted via the Connectivity Map. In summation, this study unveiled two clusters with unique molecular phenotypes within rosacea, leading to the development of a precise diagnostic model instrumental in NR diagnosis.

Humans

A regulatory network underlying idiopathic pulmonary fibrosis.

BACKGROUND: Idiopathic pulmonary fibrosis (IPF) is a progressive interstitial lung disease in which genetic susceptibility interacts with epithelial, immune, and mesenchymal remodeling. Although the chromosome 11p15.5 locus contains established IPF susceptibility signals near MUC5B and TOLLIP, the broader regulatory architecture of this region remains incompletely resolved. METHODS: We integrated IPF genome-wide association study summary statistics with methylation, expression, and protein quantitative trait loci using summary-data-based Mendelian randomization (SMR). SMR-prioritized candidates were evaluated in independent transcriptomic and methylation cohorts and further contextualized using microRNA, transcription-factor, protein-interaction, machine-learning, single-cell, and spatial transcriptomic analyses. Fibrosis-associated expression patterns were assessed in a bleomycin-induced pulmonary fibrosis rat model. RESULTS: The analyses recovered the established MUC5B and TOLLIP signals and prioritized BRSK2 as a comparatively underexplored candidate supported by eQTL-based SMR and independent molecular evidence. The BRSK2 pQTL association did not pass the HEIDI test and was therefore not interpreted as convergent protein-level genetic evidence. Network analyses linked BRSK2 to cell-cycle, metabolic-stress, and senescence-related programs, while cross-cohort machine learning prioritized FOXA2, CDC25B, and NFE2 as informative network features. Single-cell and spatial analyses localized BRSK2 preferentially to fibroblast and myofibroblast compartments and to regions with greater histological fibrosis severity. In fibrotic rat lungs, BRSK2 expression increased, whereas FOXA2 and CDC25B decreased at the transcript and protein levels. CONCLUSIONS: These findings refine the molecular landscape of the chromosome 11p15.5 IPF susceptibility locus and prioritize BRSK2 as a candidate component of an IPF-associated profibrotic fibroblast state. Its causal contribution, direct regulatory relationships, and therapeutic tractability require targeted mechanistic validation.

Idiopathic Pulmonary Fibrosis

Federated learning for the pathogenicity annotation of genetic variants in multi-site clinical settings.

MOTIVATION: Rare diseases collectively affect 5% of the population. However, fewer than 50% of rare disease patients receive a molecular diagnosis after whole genome sequencing. Supervised machine learning is a valuable approach for the pathogenicity scoring of human genetic variants. However, existing methods are often trained on curated but limited central repositories, resulting in poor accuracy when tested on external cohorts. Yet, large collections of variants generated at hospitals and research institutions remain inaccessible to machine-learning purposes because of privacy and legal constraints. Federated learning (FL) algorithms have been recently developed enabling institutions to collaboratively train models without sharing their local datasets. RESULTS: Here, we present a proof-of-concept study evaluating the effectiveness of FL for the clinical classification of genetic variants. A comprehensive array of diverse FL strategies was assessed for coding and non-coding Single Nucleotide Variants as well as Copy Number Variants. Our results showed that federated models generally achieved comparable or superior performance to traditional centralized learning. In addition, federated models reached a robust generalization to independent sets with smaller data fractions as compared to their centralized model counterparts. Our findings support the adoption of FL to establish secure multi-institutional collaborations in human variant interpretation. AVAILABILITY AND IMPLEMENTATION: All source code required to reproduce the results presented in this article, implemented in Python, is available under the GNU General Public License v3 at https://github.com/RausellLab/FedLearnVar.

Humans

Establishment of a prognostic model based on ER stress-related cell death genes and proposing a novel combination therapy in acute myeloid leukemia.

BACKGROUND: Acute myeloid leukemia (AML) is a highly heterogeneous malignancy, presenting significant challenges in accurately predicting patient prognosis. Dysregulation of endoplasmic reticulum (ER) stress and resistance to programmed cell death (PCD) are hallmarks of AML cells. However, the prognostic significance of the interplay between ER stress and cell death pathways in AML remains largely unexplored. METHODS: We analyzed RNA sequencing and clinical data from 887 AML patients across 4 cohorts to develop an ER stress-related cell death index (ERCDI) using 10 machine-learning algorithms with 117 unique combinations. Survival and time-dependent Receiver Operating Characteristic Curve (ROC) analyses were performed to assess the model's efficacy. Clinical characteristics, the tumor immune microenvironment, and drug sensitivity differences between the high- and low-risk groups were also analyzed. The CMap database was used to identify potential therapeutic drugs. In vitro and in vivo experiments, including CCK-8, colony formation, flow cytometry, Transwell assays, and xenograft mouse models, were conducted to evaluate the effects of the target genes and candidate drugs. RESULTS: The ERCDI demonstrated strong prognostic and predictive performance for prognosis in AML patients. Furthermore, the ERCDI effectively predicted immunotherapy and chemotherapy outcomes and was associated with the immune features of the different risk groups. DNA damage-inducible transcript 4 protein (DDIT4), a key gene associated with ERCDI, is related to poor prognosis in AML patients with high expression. Additionally, the knockdown of DDIT4 significantly inhibited AML cell proliferation, induced cell apoptosis, and promoted cell cycle arrest. Chaetocin was subsequently identified as a candidate compound for AML treatment. Subsequent experiments suggested that combining chaetocin and venetoclax is a potentially promising therapeutic strategy for AML. CONCLUSION: The ERCDI provides personalized risk assessment and treatment recommendations for individual AML patients. The combined use of chaetocin and venetoclax can potentially be repurposed for AML therapy.

Humans

Integration of single cell multiomics data by deep transfer hypergraph neural network.

Multi-omics characterization of individual cells offers remarkable potential for analyzing the dynamics and relationships of gene regulatory states across millions of cells. How to integrate multimodal data is an open problem, existing integration methods struggle with accuracy and modality-specific biological variation retention. In this paper, we present scHyper (scalable, interpretable machine learning for single cell integration), a low-code and data-efficient deep transfer model designed for integrating paired and unpaired single-cell multimodal data. We benchmark scHyper against datasets from different multimodal data. ScHyper learns a low-dimensional representation and aligns the covariance matrices of the measured modalities, achieving high accuracy even with large scale atlas-level datasets with low memory and computational time across different cell lines, shedding light on regulatory relationships between different types of omics. Altogether, we show that scHyper is a versatile and robust tool for cell-type label transfer and integration from multimodal single-cell datasets.

Single-Cell Analysis

Opportunities for machine learning to predict cross-neutralization in FMDV serotype O.

Accurately estimating cross-neutralization between serotype O foot-and-mouth disease viruses (FMDVs) is critical for guiding vaccine selection and disease management. In this study, we developed a machine learning approach to estimate r1 values-an established measure of antigenic similarity-using VP1 sequence data and published virus neutralization titer (VNT) results. Our dataset comprised 108 serum-virus pairs representing 73 distinct FMDV strains. We applied Boruta feature selection and random forest classifiers, optimizing model performance through tenfold cross-validation and sub-sampling to address class imbalance. Predictors included pairwise amino acid distances, site-specific polymorphisms, and differences in potential N-glycosylation sites. Using a 0.3 r1 threshold to define cross-neutralization, the final model achieved high accuracy (0.96), sensitivity (0.93), and specificity (0.96) in training, and performed robustly on independent test sets - accuracy was 0.75 (95% CI 0.60 and 0.90), F1 score 0.86% and PPV 0.77. Importantly, key VP1 residues-positions 48, 100, 135, 150, and 151-emerged as strong predictors of antigenic relationships. Our results demonstrate the utility of integrating routinely generated genomic data with machine learning to inform vaccine candidate selection and anticipate immune interactions among circulating FMDV strains. This approach offers a practical tool for accelerating vaccine decision-making and can be adapted to other FMDV serotypes. The latest version of the r1 predictive model is available for access via a Shiny dashboard (https://dmakau.shinyapps.io/PredImmune-FMD/).

Foot-and-Mouth Disease Virus

Massively parallel approaches for characterizing noncoding functional variation in human evolution.

The genetic differences underlying unique phenotypes in humans compared to our closest primate relatives have long remained a mystery. Similarly, the genetic basis of adaptations between human groups during our expansion across the globe is poorly characterized. Uncovering the downstream phenotypic consequences of these genetic variants has been difficult, as a substantial portion lies in noncoding regions, such as cis-regulatory elements (CREs). Here, we review recent high-throughput approaches to measure the functions of CREs and the impact of variation within them. CRISPR screens can directly perturb CREs in the genome to understand downstream impacts on gene expression and phenotypes, while massively parallel reporter assays can decipher the regulatory impact of sequence variants. Machine learning has begun to be able to predict regulatory function from sequence alone, further scaling our ability to characterize genome function. Applying these tools across diverse phenotypes, model systems, and ancestries is beginning to revolutionize our understanding of noncoding variation underlying human evolution.

Humans