Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “genome binning”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

VirBinn improves viral genome binning from metagenomic Hi-C through graph diffusion.

MOTIVATION: Metagenomic Hi-C provides in situ proximity signals that can improve genome binning and enable virus-host-association analysis. However, viral genome recovery remains difficult because virus-virus Hi-C contact matrices are extremely sparse. Viral genomes are small, often low-abundance, and frequently assemble into short contigs, leaving many true within-genome links unobserved and causing viral bins to fragment. RESULTS: We present VirBinn, a graph-diffusion framework for viral binning from metagenomic Hi-C. VirBinn enhances virus-virus connectivity through two complementary mechanisms: random-walk-with-restart enhancement on the sparse virus-virus contact graph and host-guided diffusion that propagates viral seeds through the host network to infer indirect virus-virus associations. The enhanced views are integrated and clustered using Leiden community detection to produce viral metagenome-assembled genomes (vMAGs). On dataset-specific simulation benchmarks with ground truth, VirBinn consistently recovers more high-quality vMAGs than Hi-C-based and shotgun-based baselines and substantially increases the number of near-complete genomes. On four real metagenomic Hi-C datasets spanning human gut, pig gut, sheep gut (long-read assembly), and wastewater, VirBinn yields more high-completeness vMAGs under CheckV and produces bins with strong within-cluster contact support. Finally, host linkage analysis using reconstructed host MAGs reveals habitat-specific host-association patterns and plausible host taxonomic profiles. AVAILABILITY AND IMPLEMENTATION: VirBinn is available at https://github.com/dyxstat/VirBinn. The scripts to reproduce the results and figures in this article are available at https://github.com/dyxstat/Reproduce_VirBinn.

Genome, Viral↗

A Graph Contrastive Learning Method for Enhancing Genome Recovery in Complex Microbial Communities.

Accurate genome binning is essential for resolving microbial community structure and functional potential from metagenomic data. However, existing approaches-primarily reliant on tetranucleotide frequency (TNF) and abundance profiles-often perform sub-optimally in the face of complex community compositions, low-abundance taxa, and long-read sequencing datasets. To address these limitations, we present MBGCCA, a novel metagenomic binning framework that synergistically integrates graph neural networks (GNNs), contrastive learning, and information-theoretic regularization to enhance binning accuracy, robustness, and biological coherence. MBGCCA operates in two stages: (1) multimodal information integration, where TNF and abundance profiles are fused via a deep neural network trained using a multi-view contrastive loss, and (2) self-supervised graph representation learning, which leverages assembly graph topology to refine contig embeddings. The contrastive learning objective follows the InfoMax principle by maximizing mutual information across augmented views and modalities, encouraging the model to extract globally consistent and high-information representations. By aligning perturbed graph views while preserving topological structure, MBGCCA effectively captures both global genomic characteristics and local contig relationships. Comprehensive evaluations using both synthetic and real-world datasets-including wastewater and soil microbiomes-demonstrate that MBGCCA consistently outperforms state-of-the-art binning methods, particularly in challenging scenarios marked by sparse data and high community complexity. These results highlight the value of entropy-aware, topology-preserving learning for advancing metagenomic genome reconstruction.

canonical correlation analysis↗

Meta-analysis of four rheumatoid arthritis genome-wide linkage studies: confirmation of a susceptibility locus on chromosome 16.

OBJECTIVE: Susceptibility to rheumatoid arthritis (RA) is likely to involve several genes of weak effect, and consequently, individual studies may have insufficient power to detect linkage. Four major RA genome-wide linkage studies have been carried out, but apart from the well-established HLA susceptibility locus, none of the reported significant regions of linkage has been replicated. We applied a genome-search meta-analysis to 4 RA genome searches to assess linkage across studies, using published results. METHODS: For each study, 120 genomic bins of approximately 30 cM were defined and ranked according to the maximum evidence for linkage within each bin. Ranks were summed across studies and each bin was assessed empirically by the magnitude of summed rank, using a permutation test. A high summed rank indicated a region in which evidence for linkage was consistent across several studies. RESULTS: In addition to the HLA locus (P < 0.00002), the strongest evidence for an RA susceptibility locus was found on chromosome 16 (P = 0.004). This locus was not identified as statistically significant in any of the 4 individual RA genome searches. In total, 12 regions achieved a significant (P < 0.05) summed rank, compared with the 6 bins expected by random chance. Four of these regions (on chromosomes 6p, 16cen, 6q, and 12p) reached a significance value of P < 0.01, suggesting that a subset of these regions contains RA susceptibility loci. CONCLUSION: Using a meta-analysis approach, we have identified existing and novel putative RA susceptibility loci. These results can provide a basis for further positional and functional candidate-gene studies, and may prove useful in other complex rheumatic diseases.

Arthritis, Rheumatoid↗

Soffritto: a deep learning model for predicting high-resolution replication timing.

MOTIVATION: Replication timing (RT) refers to the order in which DNA loci are replicated during S phase. RT is cell-type specific and implicated in cellular processes including transcription, differentiation, and disease. RT is typically quantified genome-wide using two-fraction assays (e.g. Repli-Seq) which sort cells into early and late S phase fractions followed by DNA sequencing, yielding a ratio as the RT signal. While two-fraction RT data are widely available in multiple cell lines, it is limited in its ability to capture high-resolution RT features. To address this, high-resolution Repli-Seq, which quantifies RT across 16 fractions, was developed, but it is costly and technically challenging with very limited data generated to date. RESULTS: Here, we developed Soffritto, a deep learning model that predicts high-resolution RT data using two-fraction RT data, histone ChIP-seq data, GC content, and gene density as input. Soffritto is composed of a Long Short-Term Memory (LSTM) module and a prediction module. The LSTM module learns long- and short-range interactions between genomic bins, while the prediction module is composed of a fully connected layer that outputs a 16-fraction probability vector for each bin using the LSTM module's embeddings as input. By performing both within cell line and cross-cell line training and testing for five human and mouse cell lines, we show that Soffritto is able to capture experimental 16-fraction RT signals with high accuracy, and the predicted signals allow detection of high-resolution RT patterns. AVAILABILITY AND IMPLEMENTATION: Soffritto is available at https://github.com/ay-lab/Soffritto.

Deep Learning↗

CDACHIE: chromatin domain annotation by integrating chromatin interaction and epigenomic data with contrastive learning.

MOTIVATION: Chromatin domain annotation identifies functional genomic regions, such as active and inactive zones, based on epigenomic features like histone modifications, DNA methylation, and chromatin accessibility. While recent methods have utilized both chromatin interaction data (e.g. Hi-C) and epigenomic data, they often overlook the direct relationship between these data types. RESULTS: In this study, we introduce Chromatin Domain Annotation using Contrastive Learning for Hi-C and Epigenomic Data (CDACHIE), a method for identifying chromatin domains from Hi-C and epigenomic data. Our approach leverages contrastive learning to generate aligned representative vectors for both data types at each genomic bin. The concatenated vectors are then clustered using K-means to classify distinct chromatin domain types. CDACHIE achieves superior performance in Variance Explained, evaluated across gene expression, replication timing, and ChIA-PET data. This highlights its robust ability to integrate semantic associations between Hi-C and epigenomic features within the embedding space. AVAILABILITY AND IMPLEMENTATION: The source code is available at GitHub: https://github.com/maruyama-lab-design/CDACHIE. An archival snapshot of the code used in this study is available on Zenodo: https://doi.org/10.5281/zenodo.15751780.

Chromatin↗

Enhancing genome recovery across metagenomic samples using MAGmax.

SUMMARY: The number of metagenome-assembled genomes (MAGs) is rapidly increasing with the growing scale of metagenomic studies, driving fast progress in microbiome research. Sample-wise assembly has become the standard due to its computational efficiency and strain-level resolution. It requires dereplication, the removal of near-identical genomes assembled in different metagenomic samples. We present MAGmax, an efficient dereplication tool that enhances both the quantity and quality of MAGs through a strategy of bin merging and reassembly. Unlike dRep, which selects a single representative bin per genome cluster, MAGmax merges multiple bins within a cluster and reassembles them to increase coverage. MAGmax produces more dereplicated, higher-quality MAGs than dRep at 1.6&#xd7; its speed and using three times less memory. AVAILABILITY AND IMPLEMENTATION: The MAGmax open source software, implemented in Rust, is available under the GPLv3 license at https://github.com/soedinglab/MAGmax.

Metagenomics↗

Cleanifier: contamination removal from microbial sequences using spaced seeds of a human pangenome index.

MOTIVATION: The first step when working with DNA data of human-derived microbiomes is to remove human contamination for two reasons. First, many countries have strict privacy and data protection guidelines for human sequence data, so microbiome data containing partly human data cannot be easily further processed or published. Second, human contamination may cause problems in downstream analysis, such as metagenomic binning or genome assembly. For large-scale metagenomics projects, fast and accurate removal of human contamination is therefore critical. RESULTS: We introduce Cleanifier, a fast and memory frugal alignment-free tool for detecting and removing human contamination based on gapped k-mers, or spaced seeds. Cleanifier uses a pangenome index of known human gapped k-mers, and the creation and use of alternative references is also possible. Reads are classified and filtered according to their gapped k-mer content. Cleanifier supports two filtering modes: one that queries all gapped k-mers and one that queries only a sample of them. A comparison of Cleanifier with other state-of-the-art tools shows that the sampling mode makes Cleanifier the fastest method with comparable accuracy. When using a probabilistic Cuckoo filter to store the complete k-mer set, Cleanifier has similar memory requirements to methods that use a sampled minimizer index. At the same time, Cleanifier is more flexible, because it can use different sampling methods on the same index. AVAILABILITY AND IMPLEMENTATION: Cleanifier is available via gitlab (https://gitlab.com/rahmannlab/cleanifier), PyPi (https://pypi.org/project/cleanifier/), and Bioconda (https://anaconda.org/bioconda/cleanifier). The pre-computed human pangenome index is available at Zenodo (https://doi.org/10.5281/zenodo.15639519).

Humans↗

Exploring the hypothetical role of Bacteroides species in depression progression: insights from metagenomic analysis.

Depression, a psychiatric disorder with significant morbidity and mortality, has a complex etiology. Recent advances in microbiome research have highlighted the potential role of fecal microbiota in depression pathogenesis. This study utilized shotgun metagenomic sequencing to compare the fecal microbiota of 28 depression patients and 26 healthy individuals. Significant differences in fecal microbiota composition were observed between the two groups. We generated 350 non-redundant high-quality metagenome-assembled genomes (MAGs) by binning and conducted comparisons between the depression and control groups. Notably, we found that the MAGs enriched in people with depression mostly belonged to Bacteroides, indicating a close link between Bacteroides abundance and the development of depression, suggesting that Bacteroides might be a potential culprit for depression. In the depression group, we found that the module of nitric oxide synthesis was remarkably enriched, and all Bacteroides MAGs contained genes annotated as nitric oxide synthase, suggesting that increased levels of Bacteroides may contribute to elevated nitric oxide synthesis. A distinct microbial signature consisting of Arthrobacter sp._U41, Bacillus cereus, Campylobacter rectus, and Pasteurella dagmatis accurately discriminates between depressed individuals and healthy controls, achieving an average area under the receiver operating characteristic curve of 0.950. This research sheds light on the potential role of fecal microbiota in depression and highlights specific metabolic pathways and microbial markers for further investigation.IMPORTANCEThis research highlighted significant differences in the composition and function of fecal microbiota between individuals with depression and healthy individuals, particularly the enrichment of Bacteroides metagenome-assembled genomes (MAGs) in depression patients. The upregulation of the nitric oxide synthesis pathway associated with these MAGs belonging to Bacteroides in the gut of depression patients had also been observed. The selected bacterial biomarkers reliably differentiate depression cases from healthy controls with high diagnostic accuracy (mean area under the receiver operating characteristic curve = 0.950). Our results suggest the importance of exploring microbial markers as potential diagnostic and therapeutic targets in managing depression.

Humans↗

MetaflowX: a scalable and resource-efficient workflow for multi-strategy metagenomic analysis.

Microbiomes play crucial roles in diverse ecosystems, spanning environmental, agricultural, and human health domains. However, in-depth metagenomic data analysis presents significant technical and resource challenges, particularly at scale. Existing computational pipelines are typically limited to either reference-based or reference-free approaches and exhibit inefficiencies in process large datasets. Here, we introduce MetaflowX (https://github.com/01life/MetaflowX), an open-resource workflow integrating both analytical paradigms for enhanced metagenomic investigations. This modular framework encompasses short-read quality control, rapid microbial profiling, hybrid contig assembly and binning, high-quality metagenome-assembled genome (MAG) identification, as well as bin refinement and reassembly. Benchmarking tests showed that MetaflowX completed full metagenomic analyses up to 14-fold faster and with 38% less disk usage than existing workflows. It also recovered the highest number of high-quality and taxonomically diverse MAGs. A dedicated reassembly module further improved MAG quality, increasing completeness by 5.6% and reducing contamination by 53% on average. Functional annotation modules enable detection of key features, including virulence and antibiotic resistance genes. Designed for extensibility, MetaflowX provides an efficient solution addressing current and emerging demands in large-scale metagenomic research.

Metagenomics↗

CoverM: read alignment statistics for metagenomics.

SUMMARY: Genome-centric analysis of metagenomic samples is a powerful method for understanding the function of microbial communities. Calculating read coverage is a central part of analysis, enabling differential coverage binning for recovery of genomes and estimation of microbial community composition. Coverage is determined by processing read alignments to reference sequences of either contigs or genomes. Per-reference coverage is typically calculated in an ad-hoc manner, with each software package providing its own implementation and specific definition of coverage. Here we present a unified software package CoverM which calculates several coverage statistics for contigs and genomes in an ergonomic and flexible manner. It uses "Mosdepth arrays" for computational efficiency and avoids unnecessary I/O overhead by calculating coverage statistics from streamed read alignment results. AVAILABILITY AND IMPLEMENTATION: CoverM is free software available at https://github.com/wwood/coverm. CoverM is implemented in Rust, with Python (https://github.com/apcamargo/pycoverm) and Julia (https://github.com/JuliaBinaryWrappers/CoverM_jll.jl) interfaces.

Metabolomics↗

In silico quantitative trait locus map for atherosclerosis susceptibility in apolipoprotein E-deficient mice.

OBJECTIVE: Atherosclerosis susceptibility is a genetic trait that varies between mouse strains. The goal of this study was to use a public mouse single nucleotide polymorphism (SNP) database to define the genetic loci that are associated with this trait, without the need to perform strain intercrosses that are normally required to obtain these loci. METHODS AND RESULTS: Apolipoprotein E (apoE)-deficient mice on 6 inbred genetic backgrounds were compared for atherosclerosis lesion size in the aortic root in 2 independent studies. After normalization to the C57BL/6 strain that was used in both studies, lesion areas were found in the following rank order: DBA/2J>C57BL/6>129/SV-ter>AKR/J approximately BALB/cByJ approximately C3H/HeJ. The log lesion difference in phenotypes between each of the 15 heterologous strain pairs was determined. A mouse SNP database was then used to calculate the genetic differences between the 15 strain pairs in partially overlapping 30-cM bins across the mouse genome. Correlation analyses were preformed to analyze the genetic and phenotypic differences among the strain pairs for each genetic region. The genetic regions with the highest correlations define the in silico quantitative trait loci (QTL) associated with the atherosclerosis phenotype. Five in silico atherosclerosis QTL were identified on chromosomes 1, 10, 14, 15, and 18. The loci on chromosomes 1, 10, 14, and 18 overlap with suggestive atherosclerosis QTL identified through analyses of an F(2) cohort derived from apoE-deficient mice on the C57BL/6 and FVB/N strains. CONCLUSIONS: The 5 identified in silico QTL are candidates for further study to confirm the presence and identity of atherosclerosis susceptibility genes within these loci.

Animals↗

Synteny perturbations between wheat homoeologous chromosomes caused by locus duplications and deletions correlate with recombination rates.

Loci detected by Southern blot hybridization of 3,977 expressed sequence tag unigenes were mapped into 159 chromosome bins delineated by breakpoints of a series of overlapping deletions. These data were used to assess synteny levels along homoeologous chromosomes of the wheat A, B, and D genomes, in relation to both bin position on the centromere-telomere axis and the gradient of recombination rates along chromosome arms. Synteny level decreased with the distance of a chromosome region from the centromere. It also decreased with an increase in recombination rates along the average chromosome arm. There were twice as many unique loci in the B genome than in the A and D genomes, and synteny levels between the B genome chromosomes and the A and D genome homoeologues were lower than those between the A and D genome homoeologues. These differences among the wheat genomes were attributed to differences in the mating systems of wheat diploid ancestors. Synteny perturbations were characterized in 31 paralogous sets of loci with perturbed synteny. Both insertions and deletions of loci were detected and both preferentially occurred in high recombination regions of chromosomes.

Chromosomes, Plant↗

Meta-analysis of genome-wide scans for hypertension and blood pressure in Caucasians shows evidence of susceptibility regions on chromosomes 2 and 3.

Individual genome-wide scans of blood pressure (BP) and hypertension (HT) have shown inconsistent results. The aim of this study was to investigate whether there was any consistent evidence of linkage across multiple studies with similar ethnicity. We applied the genome-search meta-analysis method (GSMA) to nine published genome-wide scans of BP (n = 5) and HT (n = 4) from Caucasian populations. For each study, the genome was divided into 120 bins and ranked according to the maximum evidence of linkage within each bin. The ranks were summed and averaged across studies and significance levels were estimated, on the basis of a distribution function of summed ranks or permutation tests without (PU) or with (PW) a study sample size weighting factor. Chromosome 3p14.1-q12.3 showed consistent evidence of linkage to HT (PU = 0.0001 and PW = 0.0001), diastolic BP (DBP) (PU = 0.007 and PW = 0.02), HT and DBP pooled (PU = 0.00002 and PW = 0.0001) and HT and systolic BP (SBP) pooled (PU = 0.0003 and PW = 0.0005). Chromosome 2p12-q22.1 showed evidence of linkage to HT (PU = 0.003 and PW = 0.009), DBP (PU = 0.05 and PW = NS), HT and DBP pooled (PU = 0.001 and PW = 0.004) and HT and SBP pooled (PU = 0.001 and P W = 0.005). The summed ranks of the HT analysis correlated significantly with those of the DBP (r = 0.20, P = 0.03) but not with those of the SBP. Both loci showed clustering of significant bins in the analysis of HT and DBP. We conclude that modest or non-significant linkage on chromosomes 3p14.1-q12.3 and 2p12-q22.1 in each individual study translates into genome-wide significant or highly suggestive linkages to HT and DBP in our GSMA analysis.

Blood Pressure↗

Comparative DNA sequence analysis of wheat and rice genomes.

The use of DNA sequence-based comparative genomics for evolutionary studies and for transferring information from model species to crop species has revolutionized molecular genetics and crop improvement strategies. This study compared 4485 expressed sequence tags (ESTs) that were physically mapped in wheat chromosome bins, to the public rice genome sequence data from 2251 ordered BAC/PAC clones using BLAST. A rice genome view of homologous wheat genome locations based on comparative sequence analysis revealed numerous chromosomal rearrangements that will significantly complicate the use of rice as a model for cross-species transfer of information in nonconserved regions.

Chromosome Mapping↗

META-DIFF: a k-mer-based pipeline that detects differentially abundant sequences in metagenomics whole genome sequencing.

Traditional case-control metagenomic studies are constrained by their dependence on taxonomic and functional databases. Because annotation occurs before differential analysis, they are limited to known elements and keep function and taxonomy separate. Although binning strategies have emerged to reconstruct genomes and mitigate this issue, they still require an assembly step, preventing the use of all available sequencing data. Here, we introduce META-DIFF, a pipeline based on differentially abundant k-mers independently of any prior annotation. From those k-mers, it reconstructs longer sequences and provides biological context, as well as the best set of unitigs to discriminate between conditions. Across both taxonomy-centric and functionally-centric benchmarks, it showed robust performance and displayed great reproducibility. It also behaved more conservatively than did other univariate methodologies, i.e. it maintained a high precision at the expense of recall, particularly in conditions of low fold-change and limited sequencing depth. The efficacy of META-DIFF was further validated through its application to a real-world colorectal cancer dataset, which produced both confirmatory and novel results compared with those of previous publications. The pipeline is able to exploit all reads and identify differentially abundant elements, including unknown DNA, prior to annotation. With the guidelines provided, META-DIFF provides users with great exploratory power to unravel microbiome changes.

Metagenomics↗

Assembly and characterization of the first complete mitochondrial genome of Epimedium sagittatum (Sieb. et Zucc.) Maxim (Berberidaceae):an invaluable traditional Chinese medicine.

BACKGROUND: Epimedium sagittatum (Sieb. et Zucc.) Maxim is an invaluable traditional Chinese medicine plant known for its properties of tonifying kidney yang, strengthening bones and muscles, and dispelling rheumatism. The chloroplast (cp) genome of E. sagittatum have been sequenced, offering critical insights for breeding and phylogenetic research. However, the mitochondrial (mt) genome of E. sagittatum remains uncharacterized, limiting comprehensive insights into its genomic evolution. RESULTS: In this study, we assembled the first complete mt genome of E. sagittatum employing Illumina and Nanopore sequencing technology and subsequently investigated comparative analysis with its closely related species. The mt genome of E. sagittatum was assembled as a multi-branched structure with a length of 339,191&#xa0;bp, within a GC content of 46.91%. Our annotation results have shown 39 protein-coding genes (PCGs), 22 tRNA genes, three rRNA genes and four pseudogenes in the E. sagittatum mt genome. The analysis of sequence repeats has detected 79 simple sequence repeats (SSRs), 10 tandem repeats and 255 dispersed repeats in the E. sagittatum mt genome. A total of 720 C to U RNA editing sites of the 34 PCGs was predicted in E. sagittatum. The codons exhibited a strong preference for A or U bases in the E. sagittatum mt genome. The analysis of nucleotide diversity (Pi) highlighted differences in genetic variability across the tested genes, with atp9 gene exhibiting the highest genetic variation. Selection pressure analysis showed that most genes were affected by negative selection during evolution, whereas ccmB, rps10, and rps12 underwent positive selection in different plants. Additionally, a Bayesian phylogenetic tree showed that E. sagittatum was closely related to E. wushanense and E. pubescens. In total of 14 homologous fragments totaling 8,954&#xa0;bp were identified between the cp and mt genomes of E. sagittatum. CONCLUSIONS: This study presents the first assembled and annotated mt genome of E. sagittatum, which provides a valuable genetic resource for the Epimedium genus and lays the foundation for investigating the phylogenetic relationship and genetic variation of this invaluable medicinal plant.

Epimedium↗

A genetic map of candidate genes and QTLs involved in tomato fruit size and composition.

In order to screen for putative candidate genes linked to tomato fruit weight and to sugar or acid content, genes and QTLs involved in fruit size and composition were mapped. Genes were selected among EST clones in the TIGR tomato EST database (http://www.tigr.org/tdb/tgi/lgi/) or corresponded to genes preferentially expressed in the early stages of fruit development. These clones were located on the tomato map using a population of introgression lines (ILs) having one segment of Lycopersicon pennellii (LA716) in a L. esculentum (M82) background. The 75 ILs allowed the genome to be segmented into 107 bins. Sixty-three genes involved in carbon metabolism revealed 79 loci. They represented enzymes involved in the Calvin cycle, glycolysis, the TCA cycle, sugar and starch metabolism, transport, and a few other functions. In addition, seven cell-cycle-specific genes mapped into nine loci. Fourteen genes, primarily expressed during the cell division stage, and 23 genes primarily expressed during the cell expansion stage, revealed 24 and 26 loci, respectively. The fruit weight, sugars, and organic acids content of each IL was measured and several QTLs controlling these traits were mapped. Comparison between map location of QTLs and candidate gene loci indicated a few candidate genes that may influence the variation of sugar or acid contents. Furthermore, the gene/QTL locations could be compared with the loci mapped in other tomato populations.

Chromosome Mapping↗