Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “genome assembly validation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

New generation pharmacogenomic tools: a SNP linkage disequilibrium Map, validated SNP assay resource, and high-throughput instrumentation system for large-scale genetic studies.

Since public and private efforts announced the first draft of the human genome last year, researchers have reported great numbers of single nucleotide polymorphisms (SNPs). We believe that the availability of well-mapped, quality SNP markers constitutes the gateway to a revolution in genetics and personalized medicine that will lead to better diagnosis and treatment of common complex disorders. A new generation of tools and public SNP resources for pharmacogenomic and genetic studies--specifically for candidate-gene, candidate-region, and whole-genome association studies--will form part of the new scientific landscape. This will only be possible through the greater accessibility of SNP resources and superior high-throughput instrumentation-assay systems that enable affordable, highly productive large-scale genetic studies. We are contributing to this effort by developing a high-quality linkage disequilibrium SNP marker map and an accompanying set of ready-to-use, validated SNP assays across every gene in the human genome. This effort incorporates both the public sequence and SNP data sources, and Celera Genomics' human genome assembly and enormous resource ofphysically mapped SNPs (approximately 4,000,000 unique records). This article discusses our approach and methodology for designing the map, choosing quality SNPs, designing and validating these assays, and obtaining population frequency ofthe polymorphisms. We also discuss an advanced, high-performance SNP assay chemisty--a new generation of the TaqMan probe-based, 5' nuclease assay-and high-throughput instrumentation-software system for large-scale genotyping. We provide the new SNP map and validation information, validated SNP assays and reagents, and instrumentation systems as a novel resource for genetic discoveries.

Alleles↗

ClinGen recuration of hearing loss-associated genes demonstrates significant changes in gene-disease validity over time.

PURPOSE: The Clinical Genome Resource (ClinGen) Hearing Loss Gene Curation Expert Panel was assembled in 2016 and has since curated 174 gene-disease relationships (GDRs) using ClinGen's semiquantitative framework. ClinGen mandates the timely recuration of all GDRs classified as Disputed, Limited, Moderate, and Strong every 2 to 3 years. METHODS: Thirty-five GDRs met the criteria for recuration within 2 years of original curation. Previous evidence was reevaluated using the latest curation guidelines, and a comprehensive literature review was performed to obtain new evidence. Recurations were approved by the Gene Curation Expert Panel and published on the ClinGen website (www.clinicalgenome.org). RESULTS: Eight of 35 GDRs (22%) changed their classification. Two Moderate and 5 Strong GDRs were upgraded to Definitive because of new case evidence. One Strong was subsumed under another Definitive GDR after evaluation of the lumping/splitting of disease entities. Twenty-seven of 35 patients remained unchanged, with little to no new evidence reported. CONCLUSION: Genes classified as Moderate and Strong were likely to build evidence and change their classification over time, whereas Limited were unlikely to gain evidence. These findings highlight the critical role of recuration in ensuring that genetic tests and research studies incorporate the most recent evidence into their efforts.

Humans↗

TERMINUS--Telomeric End-Read Mining IN Unassembled Sequences.

UNLABELLED: TERMINUS is a set of tools to map telomeres on draft sequences of whole genome shotgun sequencing projects. It mines raw sequence reads (from a trace archive) for telomeric reads, assembles them into contigs representing individual chromosome ends and BLASTs the resulting consensus sequences against the genome assembly to identify telomere-proximal genomic contigs. Finally, it estimates the sizes of telomeric gaps and identifies clones for gap closure. TERMINUS is implemented as a set of Perl scripts that requires two sets of inputs: the NCBI Trace Archive files for a given genome project; and ancillary genome assembly information. Results are output in spreadsheets containing information that facilitates manual validation. AVAILABILITY: The TERMINUS package and supplementary information can be downloaded from http://www.genome.kbrin.uky.edu/fungi_tel/terminus/ CONTACT: farman@uky.edu.

Algorithms↗

Proteomics-based validation of genomic data: applications in colorectal cancer diagnosis.

Multiple factors are involved in the translation of functional genomic results into proteins for proteome research and target validation on tumoral tissues. In this report, genes were selected by using DNA microarrays on a panel of colorectal cancer (CRC) paired samples. A large number of up-regulated genes in colorectal cancer patients were investigated for cellular location, and those corresponding to membrane or extracellular proteins were used for a non-biased expression in Escherichia coli. We investigated different sources of cDNA clones for protein expression as well as the influence of the protein size and the different tags with respect to protein expression levels and solubility in E. coli. From 29 selected genes, 21 distinct proteins were finally expressed as soluble proteins with, at least, one different fusion protein. In addition, seven of these potential markers (ANXA3, BMP4, LCN2, SPARC, SPP1, MMP7, and MMP11) were tested for antibody production and/or validation. Six of the seven proteins (all except SPP1) were confirmed to be overexpressed in colorectal tumoral tissues by using immunoblotting and tissue microarray analysis. Although none of them could be associated to early stages of the tumor, two of them (LCN2 and MMP11) were clearly overexpressed in late Dukes' stages (B and C). This proteomic study reveals novel clues for the assembly of a robust and highly efficient high throughput system for the validation of genomic data. Moreover it illustrates the different difficulties and bottlenecks encountered for performing a quick conversion of genomic results into clinically useful proteins.

Acute-Phase Proteins↗

Comparative genomics reveals hidden biosynthetic diversity in Streptomyces spp. and metal-dependent regulatory features associated with untapped specialized metabolites.

The genus Streptomyces is one of the richest sources of bioactive natural products; however, a substantial proportion of its biosynthetic gene clusters (BGCs) remain cryptic and their metabolic products are unresolved. Advances in genome mining and computational prediction now enable comprehensive exploration of this hidden biosynthetic repertoire. In this study, whole-genome sequencing and comparative genomic analyses were performed on three three newly isolated Streptomyces strains to evaluate their specialized metabolic potential. Genome assemblies were annotated and systematically analyzed using antiSMASH, DeepBGC, GECCO, and PRISM to identify, cross-validate, and functionally characterize BGCs while predicting their associated secondary metabolite scaffolds. Taxonomic analyses based on Average Nucleotide Identity (ANI), phylogenomics, and BLAST identified the isolates as Streptomyces thinghirensis, Streptomyces novocaesareae, and Streptomyces griseorubens. Applying the consensus framework across the three Streptomyces genomes yielded 43 cryptic BGCs, lacking close similarity to reference BGCs in the MIBiG database, of which 26 were classified as HIGH, 10 as MEDIUM, and 7 as LOW confidence. Notably, numerous BGCs exhibited low abundance to characterized reference clusters, indicating a high potential for previously undescribed biosynthetic pathways and novel metabolite scaffolds. Comparative analyses further revealed strain-specific biosynthetic architectures together with putative metal-responsive regulatory systems; Fur, Zur, and Nur, which were frequently associated with specialized metabolite biosynthetic loci. Collectively, these findings demonstrate the effectiveness of integrated genome-mining strategies for prioritizing cryptic biosynthetic gene clusters and highlight the remarkable biosynthetic potential of newly identified Streptomyces isolates as a source of novel natural products.

comparative genomics↗

The comprehensive mouse radiation hybrid map densely cross-referenced to the recombination map: a tool to support the sequence assemblies.

We have developed a unique comprehensive mouse radiation hybrid (RH) map of nearly 23,000 markers integrating data from three international genome centers and over 400 independent laboratories. We have cross-referenced this map to the 0.5-cM resolution recombination-based Jackson Laboratory (TJL) backcross panel map, building a complete set of RH framework chromosome maps based on a high density of known-ordered anchor markers. We have systematically typed markers to improve coverage and resolve discrepancies, and have reanalyzed data sets as needed. The cross-linking of the RH and recombination maps has resulted in a highly accurate genome-wide map with consistent marker order. We have compared these linked framework maps to the Ensemble mouse genome sequence assembly, and show that they are a useful medium resolution tool for both validating sequence assembly and elucidating chromosome biology.

Animals↗

A contiguous 3-Mb sequence-ready map in the S3-MX region on 21q22.2 based on high- throughput nonisotopic library screenings.

Progress in complete genomic sequencing of human chromosome 21 relies on the construction of high-quality bacterial clone maps spanning large chromosomal regions. To achieve this goal, we have applied a strategy based on nonradioactive hybridizations to contig building. A contiguous sequence-ready map was constructed in the Down syndrome congenital heart disease (DS-CHD) region in 21q22.2, as a framework for large-scale genomic sequencing and positional candidate gene approach. Contig assembly was performed essentially by high throughput nonisotopic screenings of genomic libraries, prior to clone validation by (1) restriction digest fingerprinting, (2) STS analysis, (3) Southern hybridizations, and (4) FISH analysis. The contig contains a total of 50 STSs, of which 13 were newly isolated. A minimum tiling path (MTP) was subsequently defined that consists of 20 PACs, 2 BACs, and 5 cosmids covering 3 Mb between D21S3 and MX1. Gene distribution in the region includes 9 known genes (c21-LRP, WRB, SH3BGR, HMG14, PCP4, DSCAM, MX2, MX1, and TMPRSS2) and 14 new additional gene signatures consisting of cDNA selection products and ESTs. Forthcoming genomic sequence information will unravel the structural organization of potential candidate genes involved in specific features of Down syndrome pathogenesis.

Chromosome Mapping↗

Genome assembly and annotation of the parasitoid jewel wasp Nasonia oneida.

The jewel wasp, Nasonia (Hymenoptera: Pteromalidae), is a well-established model system for evolutionary genetics and host-microbial interactions. Here, we present the genome of N. oneida, a species lacking prior genomic characterization, using 10× Genomics linked-read (400× coverage), Illumina short-read (120× coverage), and transcriptome data (30× coverage). The assembled genome size is 267 Mb, comprising 4,675 scaffolds, with a scaffold N50 of 1 Mb and 98.40% Benchmarking Universal Single-Copy Orthologues (BUSCOs) completeness score. Annotation revealed 32.29% (86.46 Mb) of repetitive sequences and 14,221 protein-coding genes. Comparative genomics of N. oneida with 15 other hymenopteran species validated the presence of 5,939 gene families shared among them, including 3643 single-copy and 2296 multicopy gene families. This study provides the first de novo assembly of N. oneida, providing a significant addition to the growing repertoire of molecular tools for comparative genomics and functional studies to understand the evolution of closely related species as well as the evolution of parasitic wasps.

Animals↗

Shotgun haplotyping: a novel method for surveying allelic sequence variation.

Haplotypic sequences contain significantly more information than genotypes of genetic markers and are critical for studying disease association and genome evolution. Current methods for obtaining haplotypic sequences require the physical separation of alleles before sequencing, are time consuming and are not scaleable for large surveys of genetic variation. We have developed a novel method for acquiring haplotypic sequences from long PCR products using simple, high-throughput techniques. This method applies modified shotgun sequencing protocols to sequence both alleles concurrently, with read-pair information allowing the two alleles to be separated during sequence assembly. Although the haplotypic sequences can be assembled manually from the resultant data using pre-existing sequence assembly software, we have devised a novel heuristic algorithm to automate assembly and remove human error. We validated the approach on two long PCR products amplified from the human genome and confirmed the accuracy of our sequences against full-length clones of the same alleles. This method presents a simple high-throughput means to obtain full haplotypic sequences potentially up to 20 kb in length and is suitable for surveying genetic variation even in poorly-characterized genomes as it requires no prior information on sequence variation.

Algorithms↗

Transcriptomic and Metabolomic Profiling Identifies a Core Gene-Metabolite Axis Driving African Swine Fever Virus Replication in the Soft Tick Ornithodoros lahorensis.

African swine fever virus (ASFV) causes an incurable swine disease with nearly 100% mortality, posing a catastrophic threat to global pig production. The soft tick Ornithodoros lahorensis acts as a critical biological vector that sustains persistent ASFV replication and mediates long-distance viral transmission, yet the molecular mechanisms governing ASFV-tick interplay remain poorly understood. Here, we integrated transcriptomics and metabolomics to systematically dissect molecular changes in O.&#xa0;lahorensis across three infection stages: Uninfected control, early infection (7&#x2009;days post-infection, dpi), and late persistent infection (21 dpi). Multi-omics integration revealed that ASFV extensively remodels tick host metabolism, predominantly activating purine/pyrimidine metabolism, lipid biosynthesis, and energy metabolism. We further characterized a conserved regulatory module consisting of 12 core genes and 8 signature metabolites that collectively support ASFV genome replication and virion assembly. Three hub metabolic genes (TK1, ATP5F1B, and IMPDH) were selected for functional validation via siRNA silencing in ticks; individual gene silencing suppressed ASFV loads by 89.2%, 91.5%, and 87.8%, respectively (p&#x2009;<&#x2009;0.001***). This work represents the first comprehensive multi-omics investigation of ASFV infection in O. lahorensis. We identified tick-specific molecular targets to block vector-mediated ASFV spread and established a standardized multi-omics analytical pipeline for tick-virus interaction research. Our findings elucidate the mechanistic basis of long-term ASFV persistence in soft ticks and deliver novel actionable clues for developing vector-targeted ASF intervention strategies.

Animals↗

Quality assessment of maize assembled genomic islands (MAGIs) and large-scale experimental verification of predicted genes.

Recent sequencing efforts have targeted the gene-rich regions of the maize (Zea mays L.) genome. We report the release of an improved assembly of maize assembled genomic islands (MAGIs). The 114,173 resulting contigs have been subjected to computational and physical quality assessments. Comparisons to the sequences of maize bacterial artificial chromosomes suggest that at least 97% (160 of 165) of MAGIs are correctly assembled. Because the rates at which junction-testing PCR primers for genomic survey sequences (90-92%) amplify genomic DNA are not significantly different from those of control primers ( approximately 91%), we conclude that a very high percentage of genic MAGIs accurately reflect the structure of the maize genome. EST alignments, ab initio gene prediction, and sequence similarity searches of the MAGIs are available at the Iowa State University MAGI web site. This assembly contains 46,688 ab initio predicted genes. The expression of almost half (628 of 1,369) of a sample of the predicted genes that lack expression evidence was validated by RT-PCR. Our analyses suggest that the maize genome contains between approximately 33,000 and approximately 54,000 expressed genes. Approximately 5% (32 of 628) of the maize transcripts discovered do not have detectable paralogs among maize ESTs or detectable homologs from other species in the GenBank NR nucleotide/protein database. Analyses therefore suggest that this assembly of the maize genome contains approximately 350 previously uncharacterized expressed genes. We hypothesize that these "orphans" evolved quickly during maize evolution and/or domestication.

Chromosomes, Artificial, Bacterial↗

Mapping the whole human genome by fingerprinting yeast artificial chromosomes.

Physical mapping of the human genome has until now been envisioned through single chromosome strategies. We demonstrate that by using large insert yeast artificial chromosomes (YACs) a whole genome approach becomes feasible. YACs (22,000) of 810 kb mean size (5 genome equivalents) have been fingerprinted to obtain individual patterns of restriction fragments detected by a LINE-1 (L1) probe. More than 1000 contigs were assembled. Ten randomly chosen contigs were validated by metaphase chromosome fluorescence in situ hybridization, as well as by analyzing the inter-Alu PCR patterns of their constituent YACs. We estimate that 15% to 20% of the human genome, mainly the L1-rich regions, is already covered with contigs larger than 3 Mb.

Base Sequence↗

Fantastic microbes and where to find them: evaluating learning-by-doing outcomes in a crowdfunded metagenomics workshop.

Metagenomics offers a powerful framework for authentic, interdisciplinary learning, yet it remains underrepresented in undergraduate education due to technical and infrastructural barriers. We hypothesized that a research-based, learning-by-doing metagenomics workshop supported by accessible bioinformatics tools could enhance students' perceived skills, self-efficacy, and conceptual understanding of metagenomic analysis. To test this hypothesis, we designed and evaluated a hybrid hands-on workshop in which undergraduate and postgraduate students analyzed real environmental shotgun metagenomic datasets generated from soil samples collected during a citizen science initiative. Using the graphical workflow platform KBase, participants completed an end-to-end metagenomic analysis, from quality control and assembly to genome reconstruction, taxonomic classification, functional annotation, and scientific presentation of results. Educational outcomes were assessed through validated retrospective pre-post questionnaires, self-efficacy scales, and an open-ended conceptual understanding task. Participants showed significant increases in perceived metagenomic skills and confidence in performing metagenomic analyses, while gains in perceived learning showed a positive trend. Conceptual understanding improved across educational levels, particularly among participants with limited prior experience. Together, these findings demonstrate that authentic, data-driven metagenomics activities can effectively lower barriers to computational biology and foster meaningful learning through hands-on research experiences.

Metagenomics↗

Stolbur phytoplasma genome survey achieved using a suppression subtractive hybridization approach with high specificity.

Phytoplasmas are unculturable bacterial plant pathogens transmitted by phloem-feeding hemipteran insects. DNA of phytoplasmas is difficult to purify because of their exclusive phloem location and low abundance in plants. To overcome this constraint, suppression subtractive hybridization (SSH) was modified and used to selectively amplify DNA of the stolbur phytoplasma infecting a periwinkle plant. Plasmid libraries were constructed, and the origins of the DNA inserts were verified by hybridization and PCR screenings. After a single round of SSH, there was still a significant level of contamination with plant DNA (around 50%). However, the modified SSH, which included a second round of subtraction (double SSH), resulted in an increased phytoplasma DNA purity (97%). Results validated double SSH as an efficient way to produce a genome survey for microbial agents unavailable in culture. Assembly of 266 insert sequences revealed 181 phytoplasma genetic loci which were annotated. Comparative analysis of 113 kbp indicated that among 217 protein coding sequences, 83% were homologous to "Candidatus Phytoplasma asteris" (OY-M strain) genes, with hits widely distributed along the chromosome. Most of the stolbur-specific SSH sequences were orphan genes, with the exception of two partial coding sequences encoding proteins homologous to a mycoplasma surface protein and riboflavin kinase.

Bacterial Proteins↗

Role of RNA G-Quadruplexes in the Japanese Encephalitis Virus Genome and Their Recognition as Prospective Antiviral Targets.

G-quadruplexes (GQs) have been primarily studied in the context of cancer and neurodegenerative pathologies. However, recent research has shifted focus to their existence and functional roles in viral genomes, revealing GQ-regulated key pathways in various human pathogenic viruses. While GQ structures have been reported in the genomes of emerging and re-emerging viruses, RNA viruses have been understudied compared to DNA viruses, including notable examples such as human immunodeficiency virus-1, hepatitis C virus, Ebola virus, Nipah virus, Zika virus, and SARS-CoV-2. The flavivirus family, comprising the Japanese encephalitis virus (JEV), poses a significant global threat due to recurring outbreaks yet lacks approved antivirals. In this study, we identified and characterized eight putative G-quadruplex-forming motifs within essential genes involved in genome replication, assembly, and internalization in the host cell, conserved across different JEV isolates. The formation and stability of these motifs were validated through a multitude of biophysical and cell-based assays. The interaction and binding affinity of these motifs with the known GQ-binding ligand BRACO-19 were supported by biophysical assays, confirming the capability of these motifs to form GQ structures. Notably, BRACO-19 also exerted antiviral properties through reduction of viral replication and infectious virus titers as well as inhibition of viral protein expression, as evaluated by the cell-based assays. This comprehensive molecular characterization of G-quadruplex structures within the JEV genome highlights their potential as promising antiviral targets for intervention strategies against JEV infection through GQ-specific ligands.

G-Quadruplexes↗

Comparative genomic analysis of Streptococcus parasuis and Streptococcus suis reveals mobile element-associated enrichment of antimicrobial resistance and lack of detectable same-MGE colocalization with virulence-associated genes within stable species boundaries.

Streptococcus suis is a major porcine pathogen and a zoonotic agent that causes meningitis and septicemia in humans. Streptococcus parasuis, a recently recognized close relative, remains poorly characterized with regard to its clinical significance and genomic features. In this study, we generated a single-contig closed genome assembly with genome-wide DNA methylation profiles for S. parasuis strain A1, isolated from a diseased pig in Xinjiang, China, and complemented in silico genomic predictions with isolate-level experimental validation of antimicrobial resistance (AMR) genotypes, virulence genotypes, and phenotypic susceptibility for this reference strain. Using this high-quality genome as a reference anchor, we performed comparative genomic analyses across 195 streptococcal genomes, comprising 15 S. parasuis and 180 S. suis strains, to distinguish genome-level co-occurrence of resistance and virulence determinants from their physical colocalization on the same mobile genetic element (MGE).Species boundaries remained clearly delineated at the genomic level, with a median interspecies average nucleotide identity (ANI) of approximately 86.0%, compared with intraspecies ANI medians of 97.5% for S. parasuis and 96.2% for S. suis. Pangenome analysis identified 12,693 gene clusters, of which 1086 were core clusters, and functional annotation revealed significant differences in accessory gene repertoires between the two species. Within this stable genomic framework, S. parasuis genomes carried a higher AMR gene burden; strain A1 harbored 10 AMR genes, multiple virulence-associated genes, three genomic islands, and eight prophage regions. For strain A1, PCR validation confirmed six AMR genes and six virulence genes, and disk diffusion testing demonstrated a multidrug-resistant phenotype consistent with the genotypic profile.Among 235 predicted mobile elements, 19 harbored AMR genes and seven carried Virulence Factor Database (VFDB) homologs, but none carried both categories simultaneously. This finding reflects a lack of detectable same-MGE colocalization under the applied annotation and assembly framework; it should not be interpreted as evidence of biological physical decoupling. Under a random-placement model, the expected number of co-carrying regions was only 0.57, and the probability of observing zero co-carrying regions was P&#x202f;=&#x202f;0.55. This negative result should be interpreted with caution, given the limited number of cargo-bearing regions and the predominantly draft status of most genomes. Furthermore, the A1 genome contained multiple restriction-modification systems, showed depletion of several methylation motif families in mobile regions, and had limited CRISPR spacer matching evidence, suggesting prior exposure to the relevant sequence space. None of the genomes met our predefined criteria for whole-genome convergence.Collectively, our results support a model in which S. parasuis accumulates AMR-related genes in a modular fashion via mobile elements within stable species boundaries, with no detectable same-MGE colocalization of AMR and virulence determinants under our analytical pipeline. These findings imply that AMR surveillance strategies for this species should prioritize tracking mobile genetic elements rather than inferring wholesale genomic convergence toward S. suis.

Streptococcus suis↗

Integrated Genome Mining and Bioactivity-Guided Isolation of Antimicrobial Peptides from Bacillus amyloliquefaciens BS4.

Bacterial resistance remains a critical global health challenge, driving the continuous search for novel antimicrobial agents. Bacillus amyloliquefaciens is a recognized repository of bioactive metabolites; however, its full biosynthetic potential requires integrated genomic and experimental validation. This study characterized the antimicrobial profile of B. amyloliquefaciens BS4 through a hybrid pipeline. Genome sequencing and de novo assembly revealed a 3.9&#xa0;Mb chromosome with a G&#x2009;+&#x2009;C content of 46.14%. Functional annotation identified 3,887 coding sequences, including pathways for siderophore biosynthesis and a complete bacilysin biosynthetic cluster. BGC analysis using antiSMASH v7.1.0 and BAGEL4 identified 18 biosynthetic gene clusters, while similarity network analysis via BiG-SCAPE highlighted unique singleton BGCs, indicating untapped biosynthetic diversity. Although in silico screening via Macrel predicted two putative cationic antimicrobial peptides (AMPs), bioactivity-guided purification utilizing sequential RP-HPLC, and de novo sequencing revealed a distinct set of four active peptides. Notably, three of these sequences were identified as fragments derived from the BclA exosporium protein family, highlighting the structural proteome as a non-canonical source of antimicrobials. The purified fractions exhibited activity against M. luteus and E. coli, while displaying no significant hemolytic activity or cytotoxicity, even above the MIC values. Molecular docking further supported the interaction of these candidates with bacterial targets. Overall, this hybrid strategy effectively uncovers the antimicrobial complexity of BS4, revealing 'cryptic' peptide candidates with therapeutic potential.

Bacillus amyloliquefaciens BS4↗

Unveiling novel antimicrobial peptides from the ruminant gastrointestinal microbiomes: A deep learning-driven approach yields an anti-MRSA candidate.

INTRODUCTION: Antimicrobial peptides (AMPs) present a promising avenue to combat the growing threat of antibiotic resistance. The ruminant gastrointestinal microbiome serves as a unique ecosystem that offers untapped potential for AMP discovery. OBJECTIVES: The aims of this study are to develop an effective methodology for the identification of novel AMPs from ruminant gastrointestinal microbiomes, followed by evaluating their antimicrobial efficacy and elucidating the mechanisms underlying their activity. METHODS: We developed a deep learning-based model to identify AMP candidates from a dataset comprising 120 metagenomes and 10,373 metagenome-assembled genomes derived from the ruminant gastrointestinal tract. Both in vivo and in vitro experiments were performed to examine and validate the antimicrobial activities of the AMP candidates that were selected through bioinformatic analysis and subsequently synthesized chemically. Additionally, molecular dynamics simulations were conducted to explore the action mechanism of the most potent AMP candidate. RESULTS: The deep learning model identified 27,192 potential secretory AMP candidates. Following bioinformatic analysis, 39 candidates were synthesized and tested. Remarkably, all synthesized peptides demonstrated antimicrobial activity against Staphylococcus aureus, with 79.5% showing effectiveness against multiple pathogens. Notably, Peptide 4, which exhibited the highest antimicrobial activity against methicillin-resistant Staphylococcus aureus (MRSA), confirmed this effect in a mouse model with wound infection, exhibiting a low propensity for resistance development and minimal cytotoxicity and hemolysis towards mammalian cells. Molecular dynamics simulations provided insights into the mechanism of Peptide 4, primarily its ability to disrupt bacterial cell membranes, leading to cell death. CONCLUSION: This study highlights the power of combining deep learning with microbiome research to uncover novel therapeutic candidates, paving the way for the development of next-generation antimicrobials like Peptide 4 to combat the growing threat of MRSA would infections. It also underscores the value of utilizing ruminant microbial resources.

Animals↗