Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bioinformatics software”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 955 records · Page 53Linked to original sources

[caCORE: core architecture of bioinformation on cancer research in America].

A critical factor in the advancement of biomedical research is the ease with which data can be integrated, redistributed and analyzed both within and across domains. This paper summarizes the Biomedical Information Core Infrastructure built by National Cancer Institute Center for Bioinformatics in America (NCICB). The main product from the Core Infrastructure is caCORE--cancer Common Ontologic Reference Environment, which is the infrastructure backbone supporting data management and application development at NCICB. The paper explains the structure and function of caCORE: (1) Enterprise Vocabulary Services (EVS). They provide controlled vocabulary, dictionary and thesaurus services, and EVS produces the NCI Thesaurus and the NCI Metathesaurus; (2) The Cancer Data Standards Repository (caDSR). It provides a metadata registry for common data elements. (3) Cancer Bioinformatics Infrastructure Objects (caBIO). They provide Java, Simple Object Access Protocol and HTTP-XML application programming interfaces. The vision for caCORE is to provide a common data management framework that will support the consistency, clarity, and comparability of biomedical research data and information. In addition to providing facilities for data management and redistribution, caCORE helps solve problems of data integration. All NCICB-developed caCORE components are distributed under open-source licenses that support unrestricted usage by both non-profit and commercial entities, and caCORE has laid the foundation for a number of scientific and clinical applications. Based on it, the paper expounds caCORE-base applications simply in several NCI projects, of which one is CMAP (Cancer Molecular Analysis Project), and the other is caBIG (Cancer Biomedical Informatics Grid). In the end, the paper also gives good prospects of caCORE, and while caCORE was born out of the needs of the cancer research community, it is intended to serve as a general resource. Cancer research has historically contributed to many areas beyond tumor biology. At the same time, the paper makes some suggestions about the study at the present time on biomedical informatics in China.

Computational Biology↗

RNA 3D structure prediction: (1) assessing rna 3D structure similarity from 2D structure similarity.

Computational techniques for 3D structure prediction of proteins, the holy grail of bioinformatics, have undergone major developments in recent years, geared by international cooperation and competition with CASP (Critical Assessment of Structure Prediction Techniques) like contests to improve and refine them. Although straightforward extrapolation of these methodologies for the prediction of the 3D structures of other similarly relevant bio macromolecules may not be too compelling due mostly to the intrinsic differences in constitution, nature, and function between them, the conceptual framework underlying most of those techniques applied to the development of similar computational techniques in structural biology can lead to efficient systems for prediction of the 3D structure of other bio-macromolecules. One of them is the development of rational methodologies to model RNA 3D structures from the sequence of nucleotides composing them. In this paper we establish the fundamentals of a methodology to thread a sequence of nucleotides into a set of 3D fragments extracted from a data base expressly developed for this purpose. The technique is based on a newly implemented algorithm for extraction of 3D fragments by comparison of secondary structures of RNA. The result is a highly efficient system to produce a set of fragments from which entire RNA structure for the given nucleotide sequence can be built.

Algorithms↗

The Helmholtz Network for Bioinformatics: an integrative web portal for bioinformatics resources.

SUMMARY: The Helmholtz Network for Bioinformatics (HNB) is a joint venture of eleven German bioinformatics research groups that offers convenient access to numerous bioinformatics resources through a single web portal. The 'Guided Solution Finder' which is available through the HNB portal helps users to locate the appropriate resources to answer their queries by employing a detailed, tree-like questionnaire. Furthermore, automated complex tool cascades ('tasks'), involving resources located on different servers, have been implemented, allowing users to perform comprehensive data analyses without the requirement of further manual intervention for data transfer and re-formatting. Currently, automated cascades for the analysis of regulatory DNA segments as well as for the prediction of protein functional properties are provided. AVAILABILITY: The HNB portal is available at http://www.hnbioinfo.de

Algorithms↗

Identification and characterization of mouse Erbb2 gene in silico.

The PPP1R1B-STARD3-TCAP-PNMT-MGC9753-ERBB2-MGC14832-GRB7 locus on human chromosome 17q12 is frequently amplified in human gastric and breast cancer. We have recently identified and characterized human MGC9753 (also known as wild-type CAB2) and mouse Mgc9753. Here, we identified and characterized mouse Erbb2 gene by using bioinformatics. BLAST programs revealed that mouse AK031099 cDNA was derived from mouse Erbb2 gene. Because AK031099 cDNA showed 806 C-->A nucleotide substitution compared with mouse genome draft sequences and mouse Erbb2 ESTs, the nucleotide sequence of mouse Erbb2 cDNA was determined in silico by correcting 806 A of AK031099 cDNA to C. Nucleotide position 48-3818 of mouse Erbb2 cDNA was the coding region. Mouse Erbb2 gene, consisting of 27 exons, was located within the Ppp1r1b-Grb7 locus on the mouse chromosome 11. Mouse Erbb2 protein (1256 aa) showed 87.5% total-amino-acid identity with human ERBB2 protein, and 95.2% total-amino-acid identity with rat Erbb2 protein. Mouse Ppp1r1b-Grb7 locus and human Ppp1r1b-Grb7 locus were evolutionarily conserved in the order and the orientation of genes therein. Nucleotide and amino-acid substitution rates of Neurod2 located centromeric to the Ppp1r1b-Grb7 locus were significantly lower than others within the Ppp1r1b-Grb7 locus. This is the first report on the complete coding sequence of mouse Erbb2 gene as well as on the comprehensive comparison of Ppp1r1b-Grb7 locus within the human and mouse genomes.

Amino Acid Sequence↗

Identification and characterization of PDZRN3 and PDZRN4 genes in silico.

NUMB and NUMBL are implicated in cell fate determination through the inhibition of Notch signaling. LNX, binding to NUMB and CXADR (CAR), functions as E3 ubiquitin ligase at least for NUMB. LNX is the paralog of PDZRN1 (PDZ domain containing RING finger 1). Here, we identified two novel homologs of LNX and PDZRN1 by using bioinformatics, which were designated PDZRN3 (LNX3 or SEMCAP3) and PDZRN4 (LNX4 or SAMCAP3L), respectively. KIAA1095 cDNA (AB029018) was the representative PDZRN3 cDNA. Complete coding sequence of PDZRN4 cDNA was determined by assembling nucleotide sequences of ESTs (BF059062 and AW297403), FLJ33777 cDNA (AK091096) and IMAGE5767589 cDNA (BC040922). PDZRN4 gene, consisting of 11 exons, was found to encode two isoforms with N-terminal divergence (PDZRN4 and PDZRN4S) due to an alternative promoter. PDZRN3-CNTN3 locus at human chromosome 3p13-p12.3 and PDZRN4-CNTN1 locus at human chromosome 12q12 were paralogous regions within the human genome. PDZRN3 (1066 aa) and PDZRN4 (1036 aa) showed 59.9% total-amino-acid identity. Two bipartite nuclear localization signals (NLS) were located within the C-terminal region of PDZRN3 and PDZRN4. PR34H1 and PR34H2 domains were identified as the regions conserved among PDZRN3, PDZRN4 and Drosophila CG1783. PDZRN3 and PDZRN4 consist of RING, two PDZ, PR34H1, PR34H2 domains and two NLS, while PDZRN1 and LNX consist of RING and four PDZ domains. PDZRN family proteins were classified into the LNX-PDZRN1 subfamily and the PDZRN3-PDZRN4 subfamily. This is the first report on the PDZRN3 and PDZRN4 genes.

Alternative Splicing↗

Quantitative evaluation of recall and precision of CAT Crawler, a search engine specialized on retrieval of Critically Appraised Topics.

BACKGROUND: Critically Appraised Topics (CATs) are a useful tool that helps physicians to make clinical decisions as the healthcare moves towards the practice of Evidence-Based Medicine (EBM). The fast growing World Wide Web has provided a place for physicians to share their appraised topics online, but an increasing amount of time is needed to find a particular topic within such a rich repository. METHODS: A web-based application, namely the CAT Crawler, was developed by Singapore's Bioinformatics Institute to allow physicians to adequately access available appraised topics on the Internet. A meta-search engine, as the core component of the application, finds relevant topics following keyword input. The primary objective of the work presented here is to evaluate the quantity and quality of search results obtained from the meta-search engine of the CAT Crawler by comparing them with those obtained from two individual CAT search engines. From the CAT libraries at these two sites, all possible keywords were extracted using a keyword extractor. Of those common to both libraries, ten were randomly chosen for evaluation. All ten were submitted to the two search engines individually, and through the meta-search engine of the CAT Crawler. Search results were evaluated for relevance both by medical amateurs and professionals, and the respective recall and precision were calculated. RESULTS: While achieving an identical recall, the meta-search engine showed a precision of 77.26% (+/-14.45) compared to the individual search engines' 52.65% (+/-12.0) (p < 0.001). CONCLUSION: The results demonstrate the validity of the CAT Crawler meta-search engine approach. The improved precision due to inherent filters underlines the practical usefulness of this tool for clinicians.

Data Collection↗

BLMT: statistical sequence analysis using N-grams.

UNLABELLED: Statistical analysis of amino acid and nucleotide sequences, especially sequence alignment, is one of the most commonly performed tasks in modern molecular biology. However, for many tasks in bioinformatics, the requirement for the features in an alignment to be consecutive is restrictive and "n-grams" (aka k-tuples) have been used as features instead. N-grams are usually short nucleotide or amino acid sequences of length n, but the unit for a gram may be chosen arbitrarily. The n-gram concept is borrowed from language technologies where n-grams of words form the fundamental units in statistical language models. Despite the demonstrated utility of n-gram statistics for the biology domain, there is currently no publicly accessible generic tool for the efficient calculation of such statistics. Most sequence analysis tools will disregard matches because of the lack of statistical significance in finding short sequences. This article presents the integrated Biological Language Modeling Toolkit (BLMT) that allows efficient calculation of n-gram statistics for arbitrary sequence datasets. AVAILABILITY: BLMT can be downloaded from http://www.cs.cmu.edu/~blmt/source and installed for standalone use on any Unix platform or Unix shell emulation such as Cygwin on the Windows platform. Specific tools and usage details are described in a "readme" file. The n-gram computations carried out by the BLMT are part of a broader set of tools borrowed from language technologies and modified for statistical analysis of biological sequences; these are available at http://flan.blm.cs.cmu.edu/.

Algorithms↗

PICRUSt2-SC: an update to the reference database used for functional prediction within PICRUSt2.

SUMMARY: PICRUSt2 is a bioinformatic tool that predicts microbial functions in amplicon sequencing data using a database of annotated reference genomes. We have constructed an updated database for PICRUSt2 that has substantially increased the number of bacterial (19,493 to 26,868) and archaeal (406 to 1,002) genomes as well as the number of functional annotations present. The previous PICRUSt2 database relied on many timely and computationally intensive manual processes that made it difficult to update. We constructed a new streamlined process to allow regular upgrades to the PICRUSt2 database on an ongoing basis, and used this process to create a new database, PICRUSt2-SC (Sugar-Coated). Additionally, we have shown that this updated database contains genomes that more closely match study sequences from a range of different environments. The genomes contained in the database therefore better represent these environments and this leads to an improvement in the predicted functional annotations obtained from PICRUSt2. AVAILABILITY AND IMPLEMENTATION: PICRUSt2 source code is freely available at https://github.com/picrust/picrust2 and at https://anaconda.org/bioconda/picrust2. The latest version of PICRUSt2 at the time of writing is also archived: https://doi.org/10.5281/zenodo.15119781. The PICRUSt2-SC database comes pre-installed with PICRUSt2 from version 2.6.0 onwards. Step-by-step instructions for making the updated database are at https://github.com/picrust/picrust2/wiki/Updating-the-PICRUSt2-database. All code used for the analyses and figures in this manuscript is at https://github.com/R-Wright-1/PICRUSt2-SC_application_note and https://doi.org/10.5281/zenodo.15119770.

Software↗

A hybrid genetic-neural system for predicting protein secondary structure.

BACKGROUND: Due to the strict relation between protein function and structure, the prediction of protein 3D-structure has become one of the most important tasks in bioinformatics and proteomics. In fact, notwithstanding the increase of experimental data on protein structures available in public databases, the gap between known sequences and known tertiary structures is constantly increasing. The need for automatic methods has brought the development of several prediction and modelling tools, but a general methodology able to solve the problem has not yet been devised, and most methodologies concentrate on the simplified task of predicting secondary structure. RESULTS: In this paper we concentrate on the problem of predicting secondary structures by adopting a technology based on multiple experts. The system performs an overall processing based on two main steps: first, a "sequence-to-structure" prediction is enforced by resorting to a population of hybrid (genetic-neural) experts, and then a "structure-to-structure" prediction is performed by resorting to an artificial neural network. Experiments, performed on sequences taken from well-known protein databases, allowed to reach an accuracy of about 76%, which is comparable to those obtained by state-of-the-art predictors. CONCLUSION: The adoption of a hybrid technique, which encompasses genetic and neural technologies, has demonstrated to be a promising approach in the task of protein secondary structure prediction.

Algorithms↗

SynTReN: a generator of synthetic gene expression data for design and analysis of structure learning algorithms.

BACKGROUND: The development of algorithms to infer the structure of gene regulatory networks based on expression data is an important subject in bioinformatics research. Validation of these algorithms requires benchmark data sets for which the underlying network is known. Since experimental data sets of the appropriate size and design are usually not available, there is a clear need to generate well-characterized synthetic data sets that allow thorough testing of learning algorithms in a fast and reproducible manner. RESULTS: In this paper we describe a network generator that creates synthetic transcriptional regulatory networks and produces simulated gene expression data that approximates experimental data. Network topologies are generated by selecting subnetworks from previously described regulatory networks. Interaction kinetics are modeled by equations based on Michaelis-Menten and Hill kinetics. Our results show that the statistical properties of these topologies more closely approximate those of genuine biological networks than do those of different types of random graph models. Several user-definable parameters adjust the complexity of the resulting data set with respect to the structure learning algorithms. CONCLUSION: This network generation technique offers a valid alternative to existing methods. The topological characteristics of the generated networks more closely resemble the characteristics of real transcriptional networks. Simulation of the network scales well to large networks. The generator models different types of biological interactions and produces biologically plausible synthetic gene expression data.

Algorithms↗

Data mining of Mycobacterium tuberculosis complex genotyping results using mycobacterial interspersed repetitive units validates the clonal structure of spoligotyping-defined families.

Recently, a combination of spoligotyping and bioinformatics was proposed as a potential tool for defining major circulating clades of tuberculosis bacilli. In the present study, we attempted to validate the above mentioned classification using a new high-throughput marker, named mycobacterial interspersed repetitive units (MIRUs). Using 12 MIRU loci and spoligotyping, we performed data mining of results on clinical isolates of the Mycobacterium tuberculosis complex representative of global mycobacterial allelic diversity. Knowledge rules permitting automatic labeling of major M. tuberculosis families were defined. Using this strategy, MIRU 24 appeared to be most appropriate for classifying our dataset. The Bovis family was shown to be perfectly classified by a maximum of 3 MIRUs, followed by Africanum and East African Indian (EAI) families by 4 MIRUs, the Beijing family by 6 MIRUs, Haarlem and X families by 8 MIRUs, the T family by 9, and the Latin-American and Mediterranean (LAM) family by 10 MIRUs. Considering the hierarchy of family divergence, our results corroborate a recent suggestion that EAI is the ancestral family followed by Africanum and Bovis. On the other hand, T, X, LAM and Haarlem families appear to be of more recent evolution. These results indicate that data mining of MIRUs is a valuable new tool for analyzing the evolutionary dynamics of the M. tuberculosis complex, and for monitoring an infectious disease such as tuberculosis.

Bacterial Typing Techniques↗

Computer identification of snoRNA genes using a Mammalian Orthologous Intron Database.

Based on comparative genomics, we created a bioinformatic package for computer prediction of small nucleolar RNA (snoRNA) genes in mammalian introns. The core of our approach was the use of the Mammalian Orthologous Intron Database (MOID), which contains all known introns within the human, mouse and rat genomes. Introns from orthologous genes from these three species, that have the same position relative to the reading frame, are grouped in a special orthologous intron table. Our program SNO.pl searches for conserved snoRNA motifs within MOID and reports all cases when characteristic snoRNA-like structures are present in all three orthologous introns of human, mouse and rat sequences. Here we report an example of the SNO.pl usage for searching a particular pattern of conserved C/D-box snoRNA motifs (canonical C- and D-boxes and the 6 nt long terminal stem). In this computer analysis, we detected 57 triplets of snoRNA-like structures in three mammals. Among them were 15 triplets that represented known C/D-box snoRNA genes. Six triplets represented snoRNA genes that had only been partially characterized in the mouse genome. One case represented a novel snoRNA gene, and another three cases, putative snoRNAs. Our programs are publicly available and can be easily adapted and/or modified for searching any conserved motifs within mammalian introns.

Algorithms↗

[High throughput screening and analysis of prostate cancer-related genes through mining databases].

BACKGROUND & OBJECTIVE: Investigation of differentially expressed genes in prostate cancer tissues may help to understand the molecular mechanism of prostate cancer and provide diagnostic markers or new targets for therapy. The Cancer Genome Anatomy Project (CGAP) public database provides an unprecedented opportunity for cancer researchers to mine genes differentially expressed in cancer tissues by bioinformatic methods. This study was to explore the feasibility of incorporating the Internet-available Serial Analysis of Gene Expression (SAGE) and cDNA databases to find human prostate cancer-related genes. METHODS: SAGE digital gene expression displayer (DGED) and cDNA DGED were used to analyze differentially expressed (>5 folds) genes in malignant prostate tissues compared with normal prostate tissues. The SAGE tags were filtered by their confidence and our 3 criteria. To test the confidence of tags of virtual digital analysis for gene identification, the modified GLGI (generation of longer cDNA fragments from SAGE tags for gene identification) was used to get the cDNA 3' end downstream of 20 tags. Main functions of all candidate genes and their relations to prostate cancer were annotated. RESULTS: Fifty-three differentially expressed genes were screened out by SAGE DGED, 26 of them were up-regulated and 27 were down-regulated in prostate cancer; 28 differentially expressed genes were got by cDNA DGED, 15 of them were up-regulated and 13 were down-regulated in prostate cancer. CONCLUSION: Reasonable use of public databases by the Internet-available tools is a simple, effective approach to get cancer-related genes, and might provide useful clues for further investigation although the results require experimental validation.

DNA, Complementary↗

iHOPerator: user-scripting a personalized bioinformatics Web, starting with the iHOP website.

BACKGROUND: User-scripts are programs stored in Web browsers that can manipulate the content of websites prior to display in the browser. They provide a novel mechanism by which users can conveniently gain increased control over the content and the display of the information presented to them on the Web. As the Web is the primary medium by which scientists retrieve biological information, any improvements in the mechanisms that govern the utility or accessibility of this information may have profound effects. GreaseMonkey is a Mozilla Firefox extension that facilitates the development and deployment of user-scripts for the Firefox web-browser. We utilize this to enhance the content and the presentation of the iHOP (information Hyperlinked Over Proteins) website. RESULTS: The iHOPerator is a GreaseMonkey user-script that augments the gene-centred pages on iHOP by providing a compact, configurable visualization of the defining information for each gene and by enabling additional data, such as biochemical pathway diagrams, to be collected automatically from third party resources and displayed in the same browsing context. CONCLUSION: This open-source script provides an extension to the iHOP website, demonstrating how user-scripts can personalize and enhance the Web browsing experience in a relevant biological setting. The novel, user-driven controls over the content and the display of Web resources made possible by user-scripts, such as the iHOPerator, herald the beginning of a transition from a resource-centric to a user-centric Web experience. We believe that this transition is a necessary step in the development of Web technology that will eventually result in profound improvements in the way life scientists interact with information.

Computational Biology↗

Effect of training datasets on support vector machine prediction of protein-protein interactions.

Knowledge of protein-protein interaction is useful for elucidating protein function via the concept of 'guilt-by-association'. A statistical learning method, Support Vector Machine (SVM), has recently been explored for the prediction of protein-protein interactions using artificial shuffled sequences as hypothetical noninteracting proteins and it has shown promising results (Bock, J. R., Gough, D. A., Bioinformatics 2001, 17, 455-460). It remains unclear however, how the prediction accuracy is affected if real protein sequences are used to represent noninteracting proteins. In this work, this effect is assessed by comparison of the results derived from the use of real protein sequences with that derived from the use of shuffled sequences. The real protein sequences of hypothetical noninteracting proteins are generated from an exclusion analysis in combination with subcellular localization information of interacting proteins found in the Database of Interacting Proteins. Prediction accuracy using real protein sequences is 76.9% compared to 94.1% using artificial shuffled sequences. The discrepancy likely arises from the expected higher level of difficulty for separating two sets of real protein sequences than that for separating a set of real protein sequences from a set of artificial sequences. The use of real protein sequences for training a SVM classification system is expected to give better prediction results in practical cases. This is tested by using both SVM systems for predicting putative protein partners of a set of thioredoxin related proteins. The prediction results are consistent with observations, suggesting that real sequence is more practically useful in development of SVM classification system for facilitating protein-protein interaction prediction.

Algorithms↗

Panzea: a database and resource for molecular and functional diversity in the maize genome.

Serving as a community resource, Panzea (http://www.panzea.org) is the bioinformatics arm of the Molecular and Functional Diversity in the Maize Genome project. Maize, a classical model for genetic studies, is an important crop species and also the most diverse crop species known. On average, two randomly chosen maize lines have one single-nucleotide polymorphism every approximately 100 bp; this divergence is roughly equivalent to the differences between humans and chimpanzees. This exceptional genotypic diversity underlies the phenotypic diversity maize needs to be cultivated in a wide range of environments. The Molecular and Functional Diversity in the Maize Genome project aims to understand how selection has shaped molecular diversity in maize and then relate molecular diversity to functional phenotypic variation. The project will screen 4000 loci for the signature of selection and create a wide range of maize and maize-teosinte mapping populations. These populations will be genotyped and phenotyped, permitting high-power and high-resolution dissection of the traits and relating the molecular diversity to functional variation. Panzea provides access to the genotype, phenotype and polymorphism data produced by the project through user-friendly web-based database searches and data retrieval/visualization tools, as well as a wide variety of information and services related to maize diversity.

Chromosome Mapping↗

D-ASSIRC: distributed program for finding sequence similarities in genomes.

MOTIVATION: Locating the regions of similarity in a genome requires the availability of appropriate tools such as 'Accelerated Search for SImilar Regions in Chromosomes' (ASSIRC; Vincens et al., Bioinformatics, 14, 715-725, 1998). The aim of this paper is to present different strategies for improving this program by distributing the operations and data to multiple processing units and to assess the efficiency of the different implementations in terms of running time as a function of the number of processing units. RESULTS: The new version D-ASSIRCis based on three alternative strategies of task sharing: (1) a distributed search using the splitting of studied sequences into large overlapping subsequences (strategy ASS); (2) two distributed searches for repeated exact motifs of fixed size either managed by a central processor (strategy AGD) or locally managed by numerous processors (strategy ALD). The result is that the strategy ASSis suitable for a large number of processing units (the time was divided by a factor of 12 when the number of processing units was increased from 1 to 16) wheras the strategy ALDis better for a small set of processors (typically for four or six). The different proposed strategies are efficient for various applications in genomic research, particularly for locating similarities of nucleic sequences in large genomes. AVAILABILITY: D-ASSIRCis freely available by anonymous FTP at ftp://ftp.ens.fr/pub/molbio/dassirc.tar.gz. Sources and binaries for Solaris and Linux are included in the distribution.

Algorithms↗

MetagenomicKG: a knowledge graph for metagenomic applications.

MOTIVATION: The sheer volume and variety of genomic content within microbial communities makes metagenomics a field rich in biomedical knowledge. To traverse these complex communities and their vast unknowns, metagenomic studies often depend on distinct reference databases, such as the Genome Taxonomy Database (GTDB), the Kyoto Encyclopedia of Genes and Genomes (KEGG), and the Bacterial and Viral Bioinformatics Resource Center (BV-BRC), for various analytical purposes. These databases are crucial for the genetic and functional annotation of microbial communities. Nevertheless, the inconsistent nomenclature or identifiers of these databases present challenges for effective integration, representation, and utilization. Knowledge graphs (KGs) offer an appropriate solution by organizing biological entities from different databases to standardized identifiers, allowing their interrelations to be captured into a cohesive network regardless of the naming conventions used in each source. The graph structure not only facilitates the unveiling of hidden patterns but also enriches our biological understanding with deeper insights. Despite KGs having shown potential in various biomedical fields, their application in metagenomics remains underexplored. RESULTS: We present MetagenomicKG, a novel knowledge graph specifically tailored for metagenomic analysis. MetagenomicKG integrates taxonomic, functional, and pathogenesis-related information on the human microbiome sourced from various databases, and further connects these with existing biomedical KGs to expand the biological network. Through various case studies involving the human microbiome, we demonstrate its utility in enabling hypothesis generation regarding the relationships between microbes and diseases, generating sample-specific graph embeddings, and providing robust pathogen prediction. CODE AVAILABILITY: The source code and technical details for constructing the MetagenomicKG and reproducing all analyses are available on GitHub at https://github.com/KoslickiLab/MetagenomicKG. The data used in this manuscript, including the pre-built files and use case input data, are archived on Zenodo with DOI: 10.5281/zenodo.17546861.

Metagenomics↗