Search PubMedSearch

SEARCH · Search PubMed

Results for “prokaryotic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Characterizing the ecological niche of insertion sequences within prokaryotic genomes.

Insertion sequences (ISs) are widespread prokaryotic transposable elements, often regarded as genomic parasites that primarily cause deleterious mutations. However, they can also promote adaptive changes. These antagonistic properties make their overall impact on prokaryotic evolution difficult to grasp. Here, we address this challenge by leveraging the framework of transposon ecology to analyze IS occurrences across and within 30 499 prokaryotic genomes. Combining phylogenomics with multi-scale genomic analysis, quantitative ecology, and mathematical modeling, we provide evidence that although genomes generally provide sufficient resources for IS coexistence, universal mechanisms shape their occurrence and chromosomal distribution across genomes. These include (i) the preferential localization of ISs within highly variable and GC-heterogeneous chromosomal regions of genomic plasticity, which act as the primary reservoir of IS niches; (ii) a linear scaling between IS abundance and niche size, with an average of $5.4$ additional accessible insertion sites per IS; (iii) a dependence of IS occurrence on the presence of other ISs, suggesting a form of group behavior; (iv) the accumulation of AT-rich sequences in both coding and noncoding regions up to 100 kb around ISs, indicative of ecological isolation; and (v) the spatial partitioning of mobile genetic elements around ISs, reminiscent of ecological niche differentiation. Besides these general principles, we also uncover niche specificities associated with particular IS families, hinting at regulatory mechanisms that modulate IS activity. Altogether, this comprehensive transposon ecology approach offers new insights and avenues for understanding IS-host interactions and genome evolution, moving beyond traditional host-centric perspectives.

DNA Transposable Elements

mettannotator: a comprehensive and scalable Nextflow annotation pipeline for prokaryotic assemblies.

SUMMARY: In recent years, there has been a surge in prokaryotic genome assemblies, coming from both isolated organisms and environmental samples. These assemblies often include novel species that are poorly represented in reference databases creating a need for a tool that can annotate both well-described and novel taxa, and can run at scale. Here, we present mettannotator-a comprehensive, scalable Nextflow pipeline for prokaryotic genome annotation that identifies coding and noncoding regions, predicts protein functions, including antimicrobial resistance, and delineates gene clusters. The pipeline summarizes these results in a GFF (General Feature Format) file that can be easily utilized in downstream analysis or visualized using common genome browsers. Here, we show how it works on 200 genomes from 29 prokaryotic phyla, including isolate genomes and known and novel metagenome-assembled genomes, and present metrics on its performance in comparison to other tools. AVAILABILITY AND IMPLEMENTATION: The pipeline is written in Nextflow and Python and published under an open source Apache 2.0 licence. Instructions and source code can be accessed at https://github.com/EBI-Metagenomics/mettannotator. The pipeline is also available on WorkflowHub: https://workflowhub.eu/workflows/1069.

Software

Popcorn: prediction of short coding and noncoding genomic sequences in prokaryotes.

SUMMARY: The most challenging prokaryotic genes to identify often correspond to short ORFs (sORFs) encoding small proteins or to noncoding RNAs. RNA-seq experiments commonly evince small transcripts that do not correspond to annotated genes and are candidates for novel coding sORFs or small regulatory RNAs, but it can be difficult to accurately assess whether the numerous small transcripts are coding or not. We present Popcorn (PrOkaryotic Prediction of Coding OR Noncoding), a novel machine learning method for determining whether prokaryotic sequences are coding or noncoding. We find that Popcorn is effective in distinguishing coding from noncoding sequences, including coding sORFs and noncoding RNAs. AVAILABILITY AND IMPLEMENTATION: Freely available for use on the web at https://cs.wellesley.edu/∼btjaden/Popcorn. Source code available at https://github.com/btjaden/Popcorn and https://doi.org/10.5281/zenodo.15120075.

Open Reading Frames

Investigating cross-organism prediction of prokaryotic essential proteins using unsupervised language model and ensemble strategy.

Cross-organism prediction of essential proteins is a critical task for drug discovery and microbial engineering, yet the generalizability of existing machine learning models across diverse species remains a significant challenge. In this study, we propose DeepPEP, a large language model-based framework designed to reliably transfer essential protein annotations between distantly related organisms. Utilizing 66 curated prokaryotic datasets, we systematically evaluated DeepPEP's cross-organism performance under various conditions. Initial pairwise predictions revealed a correlation between performance and evolutionary distance; however, further investigation demonstrated that integrating training data from multiple organisms yields superior predictive power. In a benchmark scenario designed to simulate real-world applications, DeepPEP outperformed the state-of-the-art tool Geptop 2.0, showcasing a robust ability to identify species-specific essential proteins. Finally, a case study on novel genomes confirmed the model's practical effectiveness. Our results suggest that DeepPEP is a powerful strategy for prokaryotic essential protein prediction, and the rigorous evaluation framework established in this study provides a new benchmark for the field.

Large Language Models

Mechanosensitive channels dominate the minimal ion channel repertoire in prokaryotes.

The eukaryotic genomes encode hundreds of proteins that function as ion channels and transporters. Essential for sustaining life, these proteins mediate the movement of inorganic ions (e.g., K+, Na+, Cl-, and Ca2+) across the plasma membrane according to their electrochemical gradients. In multicellular organisms, a diverse array of ion channels contributes to the maintenance of the resting membrane potential, the regulation of pH, osmolarity, and cell volume, and the control of secretion, electrical excitability, and synaptic activity, among many other fundamental physiological processes. Although independent evolutionary origins have been proposed for several ion channel families, their relative hierarchical importance for cellular viability remains poorly understood. To advance our knowledge of ion channel evolutionary history, we focused on determining the minimal combination of permeabilities that allows cellular viability. To this end, we conducted a survey of representative prokaryotes with small genomes across bacterial and archaeal phyla. By focusing on the smallest genomes, our approach enabled the identification of five ion channel architectures shared among prokaryotes. Among these, non-selective mechanosensitive channels (MscS and MscL) are the most abundant, followed by potassium channels, CLC-type channels and proton channels of the MotA/TolQ/ExbB family. The conservation of the mechanosensitive protein architecture across archaeal and bacterial membranes suggests that the capacity to monitor physical membrane integrity predates the requirements for electrical communication.

Journal Article

Large language models improve annotation of prokaryotic viral proteins.

Viral genomes are poorly annotated in metagenomic samples, representing an obstacle to understanding viral diversity and function. Current annotation approaches rely on alignment-based sequence homology methods, which are limited by the paucity of characterized viral proteins and divergence among viral sequences. Here we show that protein language models can capture prokaryotic viral protein function, enabling new portions of viral sequence space to be assigned biologically meaningful labels. When applied to global ocean virome data, our classifier expanded the annotated fraction of viral protein families by 29%. Among previously unannotated sequences, we highlight the identification of an integrase defining a mobile element in marine picocyanobacteria and a capsid protein that anchors globally widespread viral elements. Furthermore, improved high-level functional annotation provides a means to characterize similarities in genomic organization among diverse viral sequences. Protein language models thus enhance remote homology detection of viral proteins, serving as a useful complement to existing approaches.

Viral Proteins

Catalog of metagenome-assembled genomes of prokaryotic communities from the Red Sea hydrothermal vents.

This study presents medium- and high-quality prokaryotic metagenome-assembled genomes (MAGs) from microbial mats and sediments at Hatiba Mons, a Red Sea hydrothermal system. We recovered 1,217 bacterial and archaeal MAGs across 75 phyla, dominated by Planctomycetota and Thermoproteota. Approximately 70% of these genomes likely represent previously uncharacterized taxa.

extreme environment

Large-scale benchmarking of prokaryotic annotation tools across thousands of species.

BACKGROUND: Genome annotation is an important step in deriving functional meaning from prokaryotic sequencing data, yet systematic evaluations guiding tool selection are lacking. We present the first large-scale investigation of four prominent open-source annotation tools (Prokka, Bakta, EggNOG-mapper, and PGAP) across 156,033 diverse genomes. This includes Escherichia coli strains for baseline performance, thousands of archaea and bacteria genomes, as well as frameshifted and metagenome-assembled genomes. RESULTS: Bakta excels in annotating high-quality bacterial genomes, while PGAP was better for archaeal genomes and challenging bacterial assemblies, including metagenome-assembled, fragmented, or contaminated samples. For Gene Ontology annotation, PGAP consistently provides broader term coverage, whereas EggNOG-mapper offers more terms per feature. CONCLUSIONS: Our findings highlight tool-specific strengths crucial for selecting optimal solutions based on genome quality, taxonomy, and origin (e.g. MAGs). This study provides an evidence-based guide for users and informs future tool development.

Molecular Sequence Annotation

Identification, characterization and classification of prokaryotic nucleoid-associated proteins.

Common throughout life is the need to compact and organize the genome. Possible mechanisms involved in this process include supercoiling, phase separation, charge neutralization, macromolecular crowding, and nucleoid-associated proteins (NAPs). NAPs are special in that they can organize the genome at multiple length scales, and thus are often considered as the architects of the genome. NAPs shape the genome by either bending DNA, wrapping DNA, bridging DNA, or forming nucleoprotein filaments on the DNA. In this mini-review, we discuss recent advancements of unique NAPs with differing architectural properties across the tree of life, including NAPs from bacteria, archaea, and viruses. To help the characterization of NAPs from the ever-increasing number of metagenomes, we recommend a set of cheap and simple in vitro biochemical assays that give unambiguous insights into the architectural properties of NAPs. Finally, we highlight and showcase the usefulness of AlphaFold in the characterization of novel NAPs.

Archaea

Cross-kingdom genomic variation in chicken gut microbiomes: insights from China's diverse local breeds.

BACKGROUND: The gut microbiome possesses substantial genetic diversity that supports microbial adaptation, but the genomic variation patterns across its prokaryotic and viral populations remain incompletely characterized. RESULTS: Through integrated metagenomic and metatranscriptomic analysis of ten indigenous chicken breeds from China, we recovered 1527 representative prokaryotic MAGs, 37,555 representative DNA viral contigs, and 1867 representative RNA viral contigs (primarily comprising Bacillota/Bacteroidota, Uroviricota, and Lenarviricota/Pisuviricota, respectively). By integrating complementary short-read and long-read metagenomics with metatranscriptomics, we identified structural variants (SVs) and single-nucleotide variants (SNVs) in these cross-kingdom genomes. Positive SV-SNV density correlations occurred consistently across all microbial groups, indicating coordinated mutational processes. DNA viruses exhibited the highest variant prevalence (86.9% SNVs, 47.7% SVs), with temperate phages accumulating significantly more variants than virulent phages. Functionally, prokaryotic variants accumulated in carbohydrate metabolism and amino acid metabolism, while viral variants demonstrated broad metabolic hijacking. Horizontal gene transfer (HGT) was characterized by a strong virus-associated signature (69.40% of 536 events) and marked by an asymmetric pattern, with phage-to-bacteria (P-to-B) flow alone constituting 37.50% of all events. Random forest analysis revealed a strong bidirectional predictive relationship between SV and SNV densities across prokaryotic, DNA viral, and RNA viral populations, suggesting coupled genomic instability. Niche breadth emerged as a major driver of SNVs across kingdoms and was positively correlated with variant density. In prokaryotes, HGT events significantly shaped variant patterns. For viruses, genomic GC content was an important factor and consistently showed a negative correlation with SNV density in both DNA and RNA viruses. CONCLUSIONS: These findings demonstrate that coordinated mutational processes and kingdom-specific intrinsic factors drive genomic variation, with viruses serving as key genetic exchange vectors in chicken gut ecosystems. Video Abstract.

Animals

Anaerobic breviate protist survival in microcosms depends on microbiome metabolic function.

Anoxic and hypoxic environments serve as habitats for diverse microorganisms, including unicellular eukaryotes (protists) and prokaryotes. To thrive in low-oxygen environments, protists and prokaryotes often establish specialized metabolic cross-feeding associations, such as syntrophy, with other microorganisms. Previous studies show that the breviate protist Lenisia limosa engages in a mutualistic association with a denitrifying Arcobacter bacterium based on hydrogen exchange. Here, we investigate if the ability to form metabolic interactions is conserved in other breviates by studying five diverse breviate microcosms and their associated bacteria. We show that five laboratory microcosms of marine breviates live with multiple hydrogen-consuming prokaryotes that are predicted to have different preferences for terminal electron acceptors using genome-resolved metagenomics. Protist growth rates vary in response to electron acceptors depending on the make-up of the prokaryotic community. We find that the metabolic capabilities of the bacteria and not their taxonomic affiliations determine protist growth and survival and present new potential protist-interacting bacteria from the Arcobacteraceae, Desulfovibrionaceae, and Terasakiella lineages. This investigation uncovers potential nitrogen and sulfur cycling pathways within these bacterial populations, hinting at their roles in syntrophic interactions with the protists via hydrogen exchange.

Anaerobiosis

Maternal contact and age-dependent succession influence the assembly of the calf rumen microbiome and virome.

Early-life colonization of the rumen is particularly important; however, the processes by which microbial and viral communities are transmitted and developed remain poorly understood. Here, we present a genome-resolved investigation of the effects of maternal contact and age-dependent succession on the calf rumen microbiome and DNA virome by comparing calves raised with or without maternal contact across early life using the metagenome-assembled genomes (MAGs) and viral operational taxonomic units (vOTUs) reconstructed from whole- and virus-like particle metagenomes. Across longitudinal samples from calves and their mothers, we identified 694 MAGs and 30,479 vOTUs, substantially expanding current genome databases and revealing extensive microbial and viral novelty. Our analyses demonstrated that both prokaryotes and DNA viruses are shared between dams and calves, with greater sharing observed in calves raised with maternal contact than in calves raised without maternal contact. Notably, viral sharing between cow-calf pairs was markedly lower compared to prokaryotes, suggesting high turnover and rapid viral diversification. Age-associated analyses further revealed coordinated shifts in prokaryotes and their viruses, with dominant genera such as Prevotella, Ruminococcus, and Fibrobacter, and their corresponding viruses increasing after day 40. These findings indicate that the early-life rumen microbiome and DNA virome undergo substantial age-dependent succession and are associated with maternal contact, providing new insights into host-microbe-virus interactions during rumen development.IMPORTANCEThis study provides one of the first genome-resolved views of DNA viral community development during early rumen colonization in calves (from 1 week to 70 days of age) and reveals how maternal contact and age influence the establishment of the calf rumen microbiome and virome. By analyzing longitudinal samples from calves raised with or without their mothers, we show that prokaryotes and their viruses undergo coordinated, age-dependent succession. Our results demonstrate that maternal separation alters the assembly of the calf rumen microbiome, highlighting the influence of maternal contact during early-life rumen development. These findings underscore the high plasticity of the early-life rumen ecosystem and suggest that early management practices, such as maternal separation, can have lasting effects on rumen development. This work provides fundamental insights into the establishment and succession of the calf rumen microbiome and DNA virome during early life and may contribute to future microbiome manipulation studies.

Animals

Soil keystone viruses are regulators of ecosystem multifunctionality.

Ecosystem multifunctionality reflects the capacity of ecosystems to simultaneously maintain multiple functions which are essential bases for human sustainable development. Whereas viruses are a major component of the soil microbiome that drive ecosystem functions across biomes, the relationships between soil viral diversity and ecosystem multifunctionality remain under-studied. To address this critical knowledge gap, we employed a combination of amplicon and metagenomic sequencing to assess prokaryotic, fungal and viral diversity, and to link viruses to putative hosts. We described the features of viruses and their potential hosts in 154 soil samples from 29 farmlands and 25 forests distributed across China. Although 4,460 and 5,207 viral populations (vOTUs) were found in the farmlands and forests respectively, the diversity of specific vOTUs rather than overall soil viral diversity was positively correlated with ecosystem multifunctionality in both ecosystem types. Furthermore, the diversity of these keystone vOTUs, despite being 10-100 times lower than prokaryotic or fungal diversity, was a better predictor of ecosystem multifunctionality and more strongly associated with the relative abundances of prokaryotic genes related to soil nutrient cycling. Gemmatimonadota and Actinobacteria dominated the host community of soil keystone viruses in the farmlands and forests respectively, but were either absent or showed a significantly lower relative abundance in that of soil non-keystone viruses. These findings provide novel insights into the regulators of ecosystem multifunctionality and have important implications for the management of ecosystem functioning.

Soil Microbiology

Synthetic transcriptional repression systems in plants.

Transcriptional repression is a fundamental regulatory mechanism that enables precise control of gene expression in response to developmental signals and environmental stimuli. Synthetic biology can leverage this process within plants to engineer programmable transgene repression systems. This review examines strategies for harnessing prokaryotic repressors in eukaryotic systems to develop synthetic repression systems in plants. These systems utilize modular promoter and repressor architectures that can be tuned through operator placement and repression-domain fusion, respectively, to adjust transcriptional regulation. Chemically dependent inducibility can also be introduced either through use of native derepression mechanisms of the prokaryotic repressors or the incorporation of ligand-binding domains. Finally, this review explores key challenges in designing synthetic repression systems, including kinetics constraints, balancing ON and OFF states, and differences between transient and transgenic expression contexts. Overall, this review highlights modular design frameworks for tunable transgene expression in plants.

Gene Expression Regulation, Plant

Conservation of antiviral systems across domains of life reveals immune genes in humans.

Deciphering the immune organization of eukaryotes is important for human health and for understanding ecosystems. The recent discovery of antiphage systems revealed that various eukaryotic immune proteins originate from prokaryotic antiphage systems. However, whether bacterial antiphage proteins can illuminate immune organization in eukaryotes remains unexplored. Here, we use a phylogeny-driven approach to uncover eukaryotic immune proteins by searching for homologs of bacterial antiphage systems. We demonstrate that proteins displaying sequence similarity with recently discovered antiphage systems are widespread in eukaryotes and maintain a role in human immunity. Two eukaryotic proteins of the anti-transposon piRNA pathway are evolutionarily linked to the antiphage system Mokosh. Additionally, human GTPases of immunity-associated proteins (GIMAPs) as well as two genes encoded in microsynteny, FHAD1 and CTRC, are respectively related to the Eleos and Lamassu prokaryotic systems and exhibit antiviral activity. Our work illustrates how comparative genomics of immune mechanisms can uncover defense genes in eukaryotes.

Humans

Giants within: a new class of microbial mobile elements.

Prokaryotes harbor a diverse spectrum of extrachromosomal elements (ECEs), which are intracellular replicons maintained independently of the primary chromosome. Historically, the ECE research field has focused on relatively small ECEs, such as plasmids. However, the advent of long-read sequencing has revealed that prokaryotes also harbor various types of giant ECEs, spanning hundreds of kilobases to over 1 Mb, that were not hitherto recognized. In this review, we describe how long-read sequencing has enabled the discovery of giant ECEs and compare the genetic architectures and functional repertoires of several recently characterized examples. The functions of most genes in these ECEs remain uncharacterized, and current computational tools frequently misclassify or overlook them. We further discuss how the discovery of these giant ECEs challenges existing classification frameworks that attempt to distinguish megaplasmids, chromids, and chromosomes. Together, these findings highlight giant ECEs as a largely unexplored layer of microbial genetics, whose characterization will have broad implications for our understanding of microbial adaptation and horizontal gene transfer.

Extrachromosomal DNA

Host-virus dynamics in anaerobic digesters facing abiotic inhibition.

Viruses play a major role in controlling the structure and dynamics of microbial communities in anaerobic digesters, ecosystems sensitive to disturbances that inhibit methane production. Here, we studied the interplay between abiotic disturbances, microbiome and virome composition, and process performance, to assess whether provirus induction can be triggered by abiotic stresses known to inhibit anaerobic digestion (ammonium, phenol and sodium chloride). We monitored viral dynamics in batch mesophilic anaerobic digesters fed with biowaste through shotgun metavirome sequencing. The diversity of both prokaryotes and viruses was high, with Clostridiales dominating the prokaryotic community and Caudoviricetes dominating the viromes. We identified 132 viral contigs and 19 host genera that were differentially abundant under disturbed conditions. No significant impact of the tested abiotic stresses on provirus induction was observed under the current experimental and analytical framework. The results were consistent with viruses exerting steady, background-level predation through a putative combination of kill-the-winner dynamics at the sub-genus level and piggyback-the-winner dynamics, rather than stress-triggered, synchronous lytic bursts. A few auxiliary metabolic genes were detected, potentially targeting carbon, sulfur and cofactor metabolism in anaerobic digestion. Temperate viruses were dominant, representing up to 71% of the viral genomes confirmed as complete across all conditions. Electron microscopy analysis revealed diverse virus-like particles, including head-tailed particles typical of Caudoviricetes, but also spherical, rod-shaped and spindle-shaped particles typical of archaeal viruses. Notably, we present a new virus family, Eurekaviridae, of spindle-shaped viruses associated with methanogenic archaea.

Anaerobiosis

The planktonic microbiome of the Great Barrier Reef.

Large genome databases have markedly improved our understanding of marine microorganisms1-5. Although these resources have focused on prokaryotes, genomes from many dominant marine lineages, such as Pelagibacter and Prochlorococcus, are conspicuously underrepresented. Here we present the Great Barrier Reef Microbial Genomes Database (GBR-MGD), comprising 5,283 prokaryotic genomes obtained from Great Barrier Reef seawater samples using Nanopore and Illumina sequencing, including a collection of high-quality genomes of underrepresented groups. We show that standard short-read assemblies miss these populations owing to a combination of strain heterogeneity and low-GC-percentage sequencing bias. The GBR-MGD also comprises 20 chromosome-level picoeukaryote and 808,585 viral genomes, including a newly described clade of marine Crassvirales. We demonstrate the utility of the GBR-MGD to identify indicator taxa that can reliably predict the effects of reef management practices, such as the establishment of marine protected zones.

Bacteria