Search PubMedSearch

SEARCH · Search PubMed

Results for “Genome-wide SNPs”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

694 records · Page 8Linked to original sources

Meta-PseU: A meta-classifier for robust prediction of RNA pseudouridine modification sites from long sequences.

BACKGROUND AND OBJECTIVES: Pseudouridine (Ψ) represents one of the most abundant and conserved RNA modifications. Ψ provides an additional hydrogen-bond donor that enhances RNA structural stability and modulates translation. It participates in diverse biological processes, including RNA-protein interactions, splicing, translational control, and stress responses. Aberrant pseudouridylation is implicated in cancer, neurodegenerative disorders, and autoimmune diseases. Despite its biological importance, experimental identification of Ψ sites remains time-consuming and costly, limiting the feasibility of transcriptome-wide profiling. Computational approaches have therefore become essential complements to experimental techniques. However, state-of-the-art machine-learning and deep-learning predictors often suffer from limited generalizability due to small training datasets. To overcome these issues, we aim at constructing new long-sequence datasets and developing a novel Ψ site predictor. METHODS: New long-sequence datasets were constructed as benchmarks for RNA Ψ-site prediction. The Ψ modification sites in RMBase 3.0 were mapped to the reference genomes across three species of human, mouse, and yeast, and the RNA sequences with a length of 201 were generated by extending the upstream and downstream from the mapped, central sites. To eliminate sequence redundancy, the sequences were clustered using CD-HIT with a 70% sequence identity threshold. We developed Meta-PseU, a logistic regression-based meta-classifier that considered 118 machine learning and deep learning classifiers. The datasets and programs are freely accessible at https://github.com/kuratahiroyuki/MetaPseU. RESULTS: By optimizing model configuration, we proposed the Meta-PseU model stacking 32 machine learning and deep learning classifiers out of 118 classifiers. Meta-PseU substantially improved model generalizability, overcoming a key limitation of existing approaches. It greatly outperformed state-of-the-art predictors and achieved increasing accuracy with increasing sequence length. CONCLUSIONS: Long-sequence datasets were newly constructed as benchmarks for RNA Ψ-site prediction. Meta-PseU offers a new framework for robust Ψ-site identification by using long sequences.

Pseudouridine

Baseline Computed Tomography Coronary Angiography and Polygenic Risk Profiles in Adults With Type 2 Diabetes: A Cross-Sectional Analysis From the VOLTAIRE Study.

AIMS: To characterise baseline clinical, anatomical, and genetic cardiovascular risk profiles in participants enrolled in the VOLTAIRE (Evaluation of Polygenic Scores and CT Imaging in Risk Factor Modification in Patients with Type 2 Diabetes) study and examine concordance across these domains. METHODS: This analysis included adults with T2D who completed baseline computed tomography coronary angiography (CTCA) and polygenic risk score (PRS) assessment prior to randomisation in the VOLTAIRE study. Coronary atherosclerosis was evaluated using coronary artery calcium (CAC) score and CTCA-derived stenosis severity. Clinical risk was assessed using the New Zealand Society for the Study of Diabetes 5-year cardiovascular risk calculator. Polygenic risk for coronary artery disease was assessed using a genome-wide PRS and categorised into tertiles. RESULTS: Among 126 participants with T2D (mean age 57.5 ± 8.7 years; 62.7% male), coronary atherosclerotic burden was highly heterogeneous: 34.9% had CAC = 0, whereas 19.8% had CAC ≥ 400. Moderate-to-severe coronary stenosis (≥ 50%) was present in 40.5% of participants overall, including 20.4% of those classified as low clinical risk. PRS distribution was variable (low 37.3%, intermediate 35.7%, high 27.0%). Overlap between anatomical, genetic, and clinical domains was limited, with only 8.7% of participants classified as high risk across all three. CONCLUSIONS: Substantial heterogeneity and limited overlap exist between anatomical, genetic, and clinical cardiovascular risk measures in T2D. These findings support a multimodal approach to risk assessment integrating imaging and genetic profiling. TRIAL REGISTRATION: https://www. CLINICALTRIALS: gov; ID: NCT07091162.

Aged

Blood Metabolomic Signatures of 1-Hour Glucose Predict Cardiometabolic Risk.

BACKGROUND: Elevated 1-hour glucose levels during an oral glucose tolerance test strongly predict type 2 diabetes (T2D) and cardiovascular disease. We investigated whether the fasting blood metabolome predicting 1-hour glucose could be a target for improving β-cell function, long-term glycemic trajectories, and reducing the risks of T2D and coronary heart disease. We also investigated whether plasma microRNAs derived from key metabolic organs regulate changes in a metabolomic risk score (MRS) for predicting 1-hour glucose. METHODS: Untargeted blood metabolomics and a frequently sampled 75-g oral glucose tolerance test were performed in participants from the OmniCarb trial (n=162). In an independent weight-loss dietary intervention trial (POUNDS Lost [Preventing Overweight Using Novel Dietary Strategies]), temporal changes in MRS and plasma microRNAs measured by genome-wide sequencing were analyzed. In addition, associations of MRS at baseline and its 10-year changes with long-term risk of incident T2D and coronary heart disease were prospectively investigated in the NHS (Nurses' Health Study). RESULTS: We created a fasting blood MRS for predicting 1-hour glucose (Pearson r=0.8) and found significant associations with half-day (diurnal) postprandial glucose excursions and insulin secretion after 5-week controlled feeding interventions varying in carbohydrate amount and glycemic index. In the POUNDS Lost trial, diet-induced changes in MRSs were related to 2-year trajectories of glucose metabolism; circulating microRNAs regulating cardiometabolic abnormalities were pivotal factors influencing these changes. In the NHS, women in the top 20% of MRS had a multivariate-adjusted relative risk of 3.80 (95% CI, 2.22-6.51) for T2D and 1.48 (95% CI, 1.04-2.12) for coronary heart disease compared with those in the lowest 20%. In addition, 10-year increases in plasma metabolites related to 1-hour glucose were linearly associated with a higher risk of T2D. CONCLUSIONS: Our findings indicate that fasting blood metabolomic signatures predicting elevated 1-hour glucose reflect disease pathophysiology and could be targets for preventing T2D and coronary heart disease.

blood glucose

Methylation profiling in CNS tumor diagnostics: a single-centre real-world experience from Central Europe.

Genome-wide DNA methylation profiling has transformed neuro-oncology by providing an objective, machine learning-based taxonomy that mitigates interobserver variability and refines the histo-molecular criteria of the current WHO classification. We evaluate the real-world diagnostic performance and clinical utility of this modality in a prospective, consecutively accrued three-year cohort of 291 central nervous system (CNS) tumors across a mixed adult-pediatric population. Successful profiling was completed in 95.9% of cases. Using the Epignostix classifier, a high-confidence diagnostic match (calibrated score [CS]&#x2009;&#x2265;&#x2009;0.84) was achieved in 70.3% of analyzable samples, while 26.5% returned lower-confidence scores (&#x2265;&#x2009;0.3 to <&#x2009;0.84) and only 3.2% remained completely unclassifiable (CS&#x2009;<&#x2009;0.3). When integrated into a comprehensive diagnostic framework, methylation profiling provided clinically useful results in 81.1% of cases, establishing diagnoses in 70 cases submitted for molecular subclassification and resolving diagnostic uncertainty or prompting major revisions in 149 histologically challenging tumors. Within truly ambiguous lesions, integration of methylome data dictated tumor grade modifications in 38.8% of cases (upgrading in 29.4% and downgrading in 9.4%), shifting patient risk stratification. Crucially, over half (52.7%) of the lower-confidence cases yielded meaningful clinical integration when supported by histomorphology and ancillary genetic or immunohistochemical markers, demonstrating that rigid score cutoffs should not dictate assay failure. Discrepant or misleading classifications occurred in 1.9%. Updating bioinformatic pipelines from version 11b4 to 12.8 rescued multiple ambiguous entries, increasing overall clinical utility to 84.1%. These findings demonstrate that integrating computational epigenomics with classical neuropathology enhances diagnostic precision, while highlighting the ongoing need for careful clinical-pathological correlation.

Central nervous system tumors

Using Organoids to Unlock the Potential of Human Torpor for Spaceflight.

PURPOSE OF REVIEW: This paper reviews the current understanding of the potential for humans to enter a state of torpor/hibernation, and discusses the possibility of inducing torpor in astronauts for long-duration space travel, including some of the physiological, technological, and ethical considerations associated with its implementation. By exploring means to induce torpor in various human organoid systems, we hope such research can provides insights to comprehensive solutions to overcome some of the major hurdles that limit the potential for human to enter a state of torpor during long-duration deep-space missions, and contribute to the ongoing efforts to make such missions more feasible and safer for astronauts. RECENT FINDINGS: On future deep space missions such as NASA's planned missions to the Moon, Mars, and near-Earth asteroids, astronauts will be continuously exposed to environments that are radically different from those on Earth, each presenting multiple logistical and physiological challenges. Beyond the well-documented physiological effects of microgravity, space travelers will encounter a complex radiation environment that may contribute to significant short- and long-term adverse effects on human physiology and increase the risk of cancer and other diseases. Besides these physical challenges, life support systems must also be designed to mitigate psychological impacts of long-term isolation and confinement - all of which collectively pose formidable engineering problems. Hibernation/torpor is a state of prolonged inactivity and metabolic depression used by a wide variety of mammals to survive periods of cold temperatures and food scarcity, including some primates and perhaps even an extinct early line of hominins that lived nearly half a million years ago. Since modern humans share common ancestry with these hominins and hibernating primates, it is likely the human genome encodes the necessary genetic information to hibernate, or at least enter the similar, more transient state of torpor. The reduced body activity, lowered metabolism, and decreased energy requirements that characterize torpor suggest that developing means of inducing such a state in astronauts could address these challenges, including providing a degree of radioprotection. SUMMARY: This review explores the potential application of human torpor as a countermeasure to address the many challenges posed by long-duration spaceflight beyond low-Earth orbit (LEO), discusses various natural hibernating model systems for studying means of inducing a torpor-like state in humans, and highlights the vast potential of using human organoids to test and validate mechanisms that govern induction and maintenance of torpor to identify the means to one day safely induce this state in astronauts to provide additional protection from the myriad stressors of spaceflight.

Astronaut Health

Longitudinal whole-genome analysis of bluetongue virus identifies conserved serotype-specific genomes and distinct genomic constellations within a Colorado sheep flock (2021-2023).

Bluetongue virus (BTV) is a segmented double-stranded RNA virus of ruminants transmitted by Culicoides spp. biting midges. Although the genome consists of ten segments, classification into serotypes is primarily based on genome segment 2. However, reassortment among genomic segments is a major driver of BTV evolution and diversity. This study used longitudinal whole-genome sequencing to characterize BTV genomes collected from 2021 to 2023 within a single sheep flock in Colorado, where multiple serotypes co-circulate. Whole-genome sequences were generated from fourteen blood samples representing four serotypes: BTV-6, -11, -13, and -17. Longitudinal sampling identified multiple BTV serotypes within individual sheep across consecutive years. Tanglegram analysis comparing segment phylogenies to the segment 2 tree demonstrated incongruent topologies across all genomic segments, suggestive of reassortment or the circulation of distinct genomic constellations. Nucleotide-level comparisons revealed high sequence homology among same-serotype samples from the same year, while the greatest genetic divergence was observed among BTV-17 genomes collected in different years. Additionally, all BTV-13 genomes contained a previously undescribed nonsynonymous substitution in segment 10 predicted to extend the encoded protein by three amino acids. Together, these findings demonstrate that highly conserved BTV genomes and distinct genomic constellations can be detected at the flock level across multiple years. This longitudinal whole-genome approach reveals the genetic complexity of endemic BTV populations, including novel variants and genomic patterns consistent with reassortment that are lost with conventional serotyped-based approaches, highlighting the need to integrate whole-genome characterization into endemic BTV monitoring programs.

Animals

Genomic characterization and pathogenicity of ruminant Listeria monocytogenes isolates in a murine oral infection model.

Listeria monocytogenes is a major foodborne pathogen; its ruminant isolates display zoonotic characteristics, causing similar clinical signs in humans, including abortion and encephalitis. However, data on whole genome sequencing and pathogenicity of ruminant L. monocytogenes isolates remain sparse. This study aimed to analyze the genotypic characteristics of L. monocytogenes isolates from ruminants with listeriosis. Furthermore, we assessed the in vivo pathogenicity of four ruminant L. monocytogenes isolates, characterized via whole-genome sequencing-based genetic clustering, in orogastrically inoculated mice. The isolate LM18 (serotype 1/2b, ST224, SL6178) had the lowest lethal dose compared to the other three isolates including previous hypervirulence type (serotype 4b, ST1, SL1) and caused secondary bacteremia in lungs, with sustained bacterial loads in the spleen and liver. Genomic (listeria pathogenicity island -1 and -3) and virulence gene (actA and llsX) mutation analyses associated with virulence suggested from well-recognized studies could not elucidate the virulence of the isolates. SSI-1, which only exists in the isolate LM18 (serotype 1/2b, ST224, SL6178), may help L. monocytogenes survive in the gastrointestinal environment, thereby affecting its virulence. Further research should investigate the role of SSI-1 in the pathogenicity of L. monocytogenes. Moreover, additional studies utilizing larger datasets of ruminant isolates are required to validate our genotypic characterization and to obtain a comprehensive picture of further genotypic differences crucial for L. monocytogenes pathogenicity.

Animals

Whole-Genome Deep Learning Predicts Chemotherapy Response in Colorectal Cancer.

Chemotherapy response in colorectal cancer (CRC) exhibits significant heterogeneity, with current clinical predictors failing to capture complex genomic determinants of resistance. We developed a hybrid deep learning framework integrating convolutional neural networks (CNNs) and bidirectional long short-term memory (BiLSTM) networks to analyze whole-genome somatic mutations, evolutionary conservation, chromatin accessibility, and 3D genome architecture in 2,546 TCGA patients. An attention mechanism identified predictive genomic regions. The model achieved an AUC of 0.92 (95% CI: 0.89-0.94) in cross-validation and 0.88 (95% CI: 0.85-0.91) in independent validation, outperforming clinical models (&#x394;AUC = +0.18, p < 0.001). Key predictors included non-coding variants in TP53, KRAS, and PIK3CA regulatory regions. Triple-positive patients (mutations in all 3 regions) had significantly worse progression-free survival (HR = 4.7, p < 0.001). Our framework enables accurate chemotherapy response prediction and reveals novel non-coding resistance mechanisms, advancing precision oncology in CRC.

Humans

Draft genome sequence of Enterococcus casseliflavus strain MBBL_MP4 isolated from healthy bovine milk.

We report the draft genome sequence of Enterococcus casseliflavus MBBL_MP4, recovered from healthy bovine milk. The 3.45-Mbp genome assembly comprises 27 contigs and indicates low pathogenic potential, with no acquired antimicrobial resistance or known virulence genes. This genome provides a valuable resource for the genomic characterization of bovine-associated E. casseliflavus.

Enterococcus casseliflavus

Complete mitochondrial genomes of eight cyclophyllidean tapeworms: genome pattern and phylogenetic analysis.

Cyclophyllidean tapeworms are widespread parasites of significant medical and veterinary importance. However, mitochondrial (mt) genomic resources for cyclophyllideans from China, particularly those recovered from wildlife hosts, remain comparatively limited. In this study, we sequenced and characterized the complete mt genomes of eight cyclophyllidean isolates collected from diverse wild and domestic hosts in China, including two Hymenolepis sp. isolates and two Raillietina sp. isolates from China, and four additional isolates of previously sequenced Taenia species. The circular mt genomes ranged from 13,387 to 14,021&#xa0;bp in length, encoding 36 typical genes with variable non-coding regions. Comparative analysis revealed highly conserved gene composition and mostly conserved mt architecture, with localized rearrangement patterns detected among the cyclophyllidean lineages examined. In particular, all sampled Taeniidae exhibited a consistent trnL1-trnS2 arrangement, whereas the examined non-Taeniidae families showed the trnS2-trnL1 arrangement, confirming and extending, across additional wildlife-associated isolates, a previously proposed family-associated gene-order marker within Cyclophyllidea. Phylogenetic analyses based on concatenated amino acid sequences of the 12 protein-coding genes placed the eight isolates within their expected families, in topologies broadly consistent with previous mitogenomic studies. These data provide additional Chinese mitogenomic references, especially for underrepresented wildlife-associated isolates, and support family-associated gene-order patterns in Cyclophyllidea.

Animals

Assembly and Characterization of the First Complete Mitochondrial Genome of Tussilago farfara L.: Insights into Biological Functions and Phylogenetic Relationships within the Asteraceae Family.

Tussilago farfara L., a member of the Asteraceae family, is an economically valuable species due to its edible and medicinal properties. To elucidate the structural characteristics, genetic mechanisms, and evolutionary pathways of the organelle genomes of T. farfara, we sequenced, assembled, and annotated its mitochondrial genome for the first time. The complete mitochondrial genome of T. farfara spans 306,024&#xa0;bp and contains 33 mitochondrial protein-coding genes (PCGs), 3 rRNAs, and 22 tRNAs. Analysis of the nucleotide substitution rate and genetic diversity revealed that most mitochondrial genome genes may have undergone purifying selection, indicating a slow evolutionary rate and a relatively conserved genomic structure. We further identified 13 fragments of chloroplast-derived DNA integrated into the mitochondrial genome, evidencing intracellular gene transfer. Collinearity analysis showed that Arctium lappa shares the most extensive mitochondrial homologous sequences and the highest sequence similarity with T. farfara. Phylogenetic analysis based on the mitochondrial genome helped to clarify the evolutionary and taxonomic position of T. farfara within the Asteraceae family. The mitochondrial genome sequence of T. farfara provides a valuable genomic resource for species identification and for evolutionary studies within the Asteraceae family.

Genome, Mitochondrial

Plant species identification by genome skimming across the vascular plant tree of life.

Accurate species identification is essential for biodiversity conservation and sustainable use, yet standard plant DNA barcoding often fails to achieve species-level resolution. We present a large-scale empirical evaluation of genome skimming as a tool to improve plant species discrimination. Using standardised data from 1969 individuals representing 475 species from 32 genera across major lineages of the vascular plant tree of life, we compare conventional plastid + internal transcribed spacer (ITS) barcodes with genome skimming approaches. Standard barcoding using rbcL, matK, trnH-psbA and ITS resolved about half of species (49.3%), with six genera showing <&#x2009;25% species discrimination. By contrast, genome skimming enabled the recovery of complete plastid genomes, yielding 57.6% species discrimination. It also generated sufficient nuclear genomic data for additional resolution from k-mer analysis, achieving 66.8% species discrimination - an average gain of 17.5% over standard barcodes - while eliminating cases of extreme failure (<&#x2009;25% resolution). The recovery of complete plastomes and ribosomal DNAs from genome skims also ensures backward compatibility with existing barcode datasets. Our results demonstrate that genome skimming provides data that substantially improves species-level resolution across diverse plant lineages and offers a scalable, high-throughput approach for building comprehensive reference resources to support global biodiversity initiatives.

DNA Barcoding, Taxonomic

AI-enabled viral genomics: from virus discovery to host prediction and emerging variant forecasting.

The rapid expansion of metagenomic sequencing has generated vast repositories of viral sequence data that far outpace our capacity to interpret them using conventional approaches. Highly divergent sequences, sparse functional annotation, and taxonomically uneven sampling present fundamental challenges for reference-dependent methods, which lose sensitivity precisely for novel and understudied viruses with high public health relevance. Artificial intelligence (AI) provides a new avenue to address these challenges by enabling predictive inference from viral genomes and proteins while reducing dependence on sequence similarity. In this Review, we discuss representative advances in AI for virus discovery, taxonomic classification and functional annotation, prediction of host range and zoonotic potential, and efforts toward forecasting emerging variants. These advances are transforming viral genomics from a largely descriptive discipline into one with increasing predictive capability. We also critically assess the major challenges that constrain current approaches, including the availability of high-quality and representative datasets, rigorous model evaluation, biological interpretability and responsible governance for increasingly capable AI models.

Artificial Intelligence

Comparative genomic epidemiology of food- and patient-derived diarrheagenic Escherichia coli from sentinel surveillance in Southeast China.

Diarrheagenic Escherichia coli (DEC) remains an important foodborne pathogen, yet long-term comparative genomic surveillance data jointly characterizing food-derived and patient-derived isolates remain limited. This surveillance-based comparative study integrated antimicrobial susceptibility testing and whole-genome sequencing to characterize diarrheagenic Escherichia coli isolates recovered from food and patient sources in Lishui, Southeast China, during 2018-2025, with emphasis on occurrence, resistance profiles, genomic backgrounds, and plasmid replicon-associated features. Antimicrobial susceptibility testing was performed for 258 selected isolates, and whole-genome sequencing was conducted for a curated analytical subset of 204 isolates. The sequenced subset was used for diversity-oriented comparative genomic analysis rather than for unbiased prevalence estimation of the entire DEC collection. EAEC predominated in both sources, although food-associated occurrence was heterogeneous across categories, with the highest recovery rate observed in raw meat. Patient-derived isolates showed a broader overall resistance burden, whereas food-derived isolates retained substantial resistance to tetracycline, chloramphenicol, and florfenicol. Phylogenetic analysis showed partial overlap in genomic backgrounds between food-derived and patient-derived isolates, while representative resistance determinants displayed both broadly distributed and lineage-enriched patterns. Replicon-based plasmid profiling identified 42 plasmid types, including 12 detected in both sources, with IncF-related replicons predominating among these shared profiles. Several food-derived isolates carried multiple plasmid replicon types that were also observed in patient-derived isolates. Overall, food-derived and patient-derived DEC showed partial overlap in genomic backgrounds, resistance determinants, and replicon-defined plasmid profiles within this surveillance setting, while retaining source-associated heterogeneity. These findings should be interpreted as surveillance-based comparative evidence rather than as evidence of direct source attribution or transmission.

Humans