Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 739 records · Page 41Linked to original sources

Species-specific monoclonal antibodies to Escherichia coli-expressed p36 cytosolic protein of Mycoplasma hyopneumoniae.

The p36 protein of Mycoplasma hyopneumoniae is a cytosolic protein carrying species-specific antigenic determinants. Based on the genomic sequence of the reference strain ATCC 25934, primers were designed for PCR amplification of the p36-encoding gene (948 bp). These primers were shown to be specific to M. hyopneumoniae since no DNA amplicons could be obtained with other mycoplasma species and pathogenic bacteria that commonly colonize the porcine respiratory tract. The amplified p36 gene was subcloned into the pGEX-4T-1 vector to be expressed in Escherichia coli as a fusion protein with glutathione S-transferase (GST). The GST-p36 recombinant fusion protein was purified by affinity chromatography and cut by thrombin, and the enriched p36 protein was used to immunize female BALB/c mice for the production of anti-p36 monoclonal antibodies (MAbs). The polypeptide specificity of the nine MAbs obtained was confirmed by Western immunoblotting with cell lysates prepared from the homologous strain. Cross-reactivity studies of the anti-p36 MAbs towards two other M. hyopneumoniae reference strains (ATCC 25095 and J strains) and Quebec field strains that had been isolated in culture suggested that these anti-p36 MAbs were directed against a highly conserved epitope, or closely located epitopes, of the p36 protein. No reactivity was demonstrated against other mycoplasma species tested. Clinical signs and lesions suggestive of enzootic pneumonia were reproduced in specific-pathogen-free pigs infected experimentally with a virulent Quebec field strain (IAF-DM9827) of M. hyopneumoniae. The bacteria could be recovered from lung homogenates of pigs that were killed after the 3-week observation period by both PCR and cultivation procedures. Furthermore, the anti-p36 MAbs permitted effective detection by indirect immunofluorescence of M. hyopneumoniae in frozen lung sections from experimentally infected pigs. However, attempts to use the recombinant p36 protein as an antigen in an indirect enzyme-linked immunosorbent assay for the detection of antibodies in sera from convalescent pigs showed no correlation with clinical and pathological findings.

Animals↗

HSP60 gene sequences as universal targets for microbial species identification: studies with coagulase-negative staphylococci.

A set of universal degenerate primers which amplified, by PCR, a 600-bp oligomer encoding a portion of the 60-kDa heat shock protein (HSP60) of both Staphylococcus aureus and Staphylococcus epidermidis were developed. However, when used as a DNA probe, the 600-bp PCR product generated from S. epidermidis failed to cross-hybridize under high-stringency conditions with the genomic DNA of S. aureus and vice versa. To investigate whether species-specific sequences might exist within the highly conserved HSP60 genes among different staphylococci, digoxigenin-labelled HSP60 probes generated by the degenerate HSP60 primers were prepared from the six most commonly isolated Staphylococcus species (S. aureus 8325-4, S. epidermidis 9759, S. haemolyticus ATCC 29970, S. schleiferi ATCC 43808, S. saprophyticus KL122, and S. lugdunensis CRSN 850412). These probes were used for dot blot hybridization with genomic DNA of 58 reference and clinical isolates of Staphylococcus and non-Staphylococcus species. These six Staphylococcus species HSP60 probes correctly identified the entire set of staphylococcal isolates. The species specificity of these HSP60 probes was further demonstrated by dot blot hybridization with PCR-amplified DNA from mixed cultures of different Staphylococcus species and by the partial DNA sequences of these probes. In addition, sequence homology searches of the NCBI BLAST databases with these partial HSP60 DNA sequences yielded the highest matching scores for both S. epidermidis and S. aureus with the corresponding species-specified probes. Finally, the HSP60 degenerate primers were shown to amplify an anticipated 600-bp PCR product from all 29 Staphylococcus species and from all but 2 of 30 other microbial species, including various gram-positive and gram-negative bacteria, mycobacteria, and fungi. These preliminary data suggest the presence of species-specific sequence variation within the highly conserved HSP60 genes of staphylococci. Further work is required to determine whether these degenerate HSP60 primers may be exploited for species-specific microbic identification and phylogenetic investigation of staphylococci and perhaps other microorganisms in general.

Base Sequence↗

Polymorphic amplified typing sequences provide a novel approach to Escherichia coli O157:H7 strain typing.

Escherichia coli O157:H7 (O157) strains are commonly typed by pulsed-field gel electrophoresis (PFGE) following digestion of genomic DNA with the restriction enzyme XbaI. We have shown that O157 strains differ from each other by a series of discrete insertions or deletions, some of which contain recognition sites for XbaI, suggesting that these insertions and deletions are responsible for the differences in PFGE patterns. We have devised a new O157 strain typing protocol, polymorphic amplified typing sequences (PATS), based on this information. We designed PCR primer pairs to amplify genomic DNA flanking each of 40 individual XbaI sites in the genomes of two O157 reference strains. These primer pairs were tested with 44 O157 isolates, 2 each from 22 different outbreaks of infection. Thirty-two primer pairs amplified identical fragments from all 44 isolates, while eight primer pairs amplified regions that were polymorphic between isolates. The isolates could be differentiated solely on the basis of which of the eight polymorphic amplicons was detected. PATS correctly identified 21 of 22 outbreak pairs as being identical or highly related, whereas PFGE correctly identified 14 of the 22 outbreak pairs as being identical or highly related; PATS was also able to type isolates from three outbreaks that were untypeable by PFGE. However, PATS was less sensitive than PFGE in discriminating between outbreaks. These data suggest that typing by PATS may provide a simple procedure for strain typing of O157 and other bacteria and that further evaluation of the utility of this method for epidemiologic investigations is warranted.

Bacterial Typing Techniques↗

Filling the gaps in the porcine linkage map: isolation of microsatellites from chromosome 18 using flow sorting and SINE-PCR.

Flow-sorted chromosome 18 material from the pig (SSC18) was amplified with SINE-PCR to generate DNA for molecular cloning. An SSC18-enriched library was constructed and subsequently screened with a (CA)15 probe to identify polymorphic microsatellites. Eleven unique microsatellites were obtained, six of which constituted direct extensions of SINE poly(A) tracts. Eight primer pairs amplified polymorphic loci in unrelated pigs. The polymorphic markers were typed in a Swedish reference pedigree for porcine genome mapping, but only two of them (S0177 and S0179) mapped to SSC18. The remaining markers mapped to chromosomes 3, 4, 7, 9, and 11. Together with two markers from other sources (S0062 and sw787), a four-point linkage group on SSC18 could be established.

Animals↗

Chimeric restriction enzymes: what is next?

Chimeric restriction enzymes are a novel class of engineered nucleases in which the non-specific DNA cleavage domain of Fokl (a type IIS restriction endonuclease) is fused to other DNA-binding motifs. The latter include the three common eukaryotic DNA-binding motifs, namely the helix-turn-helix motif, the zinc finger motif and the basic helix-loop-helix protein containing a leucine zipper motif. Such chimeric nucleases have been shown to make specific cuts in vitro very close to the expected recognition sequences. The most important chimeric nucleases are those based on zinc finger DNA-binding proteins because of their modular structure. Recently, one such chimeric nuclease, Zif-QQR-F(N) was shown to find and cleave its target in vivo. This was tested by microinjection of DNA substrates and the enzyme into frog oocytes (Carroll et al., 1999). The injected enzyme made site-specific double-strand breaks in the targets even after assembly of the DNA into chromatin. In addition, this cleavage activated the target molecules for efficient homologous recombination. Since the recognition specificity of zinc fingers can be manipulated experimentally, chimeric nucleases could be engineered so as to target a specific site within a genome. The availability of such engineered chimeric restriction enzymes should make it feasible to do genome engineering, also commonly referred to as gene therapy.

DNA Restriction Enzymes↗

Porifera a reference phylum for evolution and bioprospecting: the power of marine genomics.

The term Urmetazoa, as the hypothetical metazoan ancestor, was introduced to highlight the finding that all metazoan phyla including the Porifera [sponges] derived from one common ancestor. Analyses of sponge genomes, from Demospongiae, Calcarea and Hexactinellida have permitted the reconstruction of the evolutionary trail from Fungi to Metazoa. This has provided evidence that the characteristic evolutionary novelties of Metazoa existing in Porifera share high sequence similarities and in some aspects also functional similarities to related polypeptides found in other metazoan phyla. It is surprising that the genome of Porifera is large and comprises substantially more genes than Protostomia and Deuterostomia. On the basis of solid taxonomy and ecological data, the high value of this phylum for human application becomes obvious especially with regard to the field of chemical ecology and the hope to find novel potential drugs for clinical use. In addition, the benefit of efforts in understanding molecular biodiversity with focus on sponges can be seen in the fact that these animals as "living fossils" allow to stethoscope into the past of our globe especially with respect to the evolution of Metazoa.

Animals↗

Murine eotaxin-2: a constitutive eosinophil chemokine induced by allergen challenge and IL-4 overexpression.

The generation of tissue eosinophilia is governed in part by chemokines; initial investigation has identified three chemokines in the human genome with eosinophil selectivity, referred to as eotaxin-1, -2, and -3. Elucidation of the role of these chemokines is dependent in part upon analysis of murine homologues; however, only one murine homologue, eotaxin-1, has been identified. We now report the characterization of the murine eotaxin-2 cDNA, gene and protein. The eotaxin-2 cDNA contains an open reading frame that encodes for a 119-amino acid protein. The mature protein, which is predicted to contain 93 amino acids, is most homologous to human eotaxin-2 (59.1% identity), but is only 38.9% identical with murine eotaxin-1. Northern blot analysis reveals three predominant mRNA species and highest constitutive expression in the jejunum and spleen. Additionally, allergen challenge in the lung with Aspergillus fumigatus or OVA revealed marked induction of eotaxin-2 mRNA. Furthermore, eotaxin-2 mRNA was strongly induced by both transgenic over-expression of IL-4 in the lung and administration of intranasal IL-4. Analysis of eotaxin-2 mRNA expression in mice transgenic for IL-4 but genetically deficient in STAT-6 revealed that the IL-4-induced expression was STAT-6 dependent. Recombinant eotaxin-2 protein induced dose-dependent chemotactic responses on murine eosinophils at concentrations between 1-1000 ng/ml, whereas no activity was displayed on murine macrophages or neutrophils. Functional analysis of recombinant protein variants revealed a critical role for the amino terminus. Thus, murine eotaxin-2 is a constitutively expressed eosinophil chemokine likely to be involved in homeostatic, allergen-induced, and IL-4-associated immune responses.

Administration, Intranasal↗

Distinct rotaviruses isolated from asymptomatic calves.

Rotaviruses were isolated following cell culture of the intestinal contents of four non-diarrheic calves. The four isolates were serially propagated in MDBK and BSC-1 cells in the presence of trypsin and produced rotavirus particles morphologically similar to those found associated with diarrhea. They were antigenically related to the Nebraska calf rotavirus (Norden strain) as investigated by immunofluorescence. Three isolates could be distinguished from the reference Nebraska rotavirus by their thermal stability and/or their differential responses to intestinal neutralizing antibodies. Two isolates produced on BSC-1 cells plaques significantly different in size from the reference strain, No significant genomic variations were detected among the isolates.

Animals↗

Congenital genetic instability in colorectal carcinomas.

INTRODUCTION: Oncogenic evolution is probably based on a progressive selection of clonal subpopulations from within a single clone. This selection is supposed to be based on enhanced genetic instability in the genome. In the vast majority of patients this increased lability is supposed to be the result of acquired alterations, and once established, it may contribute to the continuing, genetic instability within the neoplastic cells. It has been postulated, that inborn chromosomal instability is not limited to a few rare syndromes. Indeed, one of the common colorectal cancer (CRC) syndromes, familial adenomatous polyposis (FAP), is supposed to be a chromosomal instability syndrome. Also the hereditary nonpolyposis colorectal cancer syndrome (Lynch Syndrome (LS)) has shown genomic instability. Knudson (1971, 1985) has shown, that the same gene can be involved in both the hereditary form of a cancer and in the sporadic form of the same cancer. Due to the existence of such hereditary chromosomal instability syndromes, inherited genetic instability may be of some importance in the evolution of sporadic colorectal cancers. METHODS: In vitro research on dermal fibroblasts is based on the theory, that all studied cells of an individual carry the same genetic material, irrespective of their in vivo expression. Increased in vitro tetraploidy (IVT+) in skin fibroblast cultures is supposed to be a germinally transmitted expression of genomic instability, with special reference to LS. At the same time the chromosomal aberrations, which occur in neoplasms, can be measured by flow cytometric DNA analysis. The DNA content, thus measured, is supposed to be an expression of somatic acquired genetic instability. Finally, since occult mandibular osteomas have been shown to be associated with FAP, we have investigated a substantial part of our patients with this phenotypical marker. AIM OF STUDY: In CRC the adenoma-carcinoma sequence is widely accepted, and the purpose of our investigation was to find a correlation, if any, between: Changes in adenoma flow cytometric DNA content and histological grade and type of adenomas. Changes in dermal fibroblast IVT+ from adenoma patients and histological grade and type of adenomas. Changes in dermal fibroblast IVT+ and flow cytometrical DNA content in adenomas and carcinomas from the same patients. Furthermore, we wanted to investigate if: The occurrence of IVT+ in skin fibroblasts among patients with CRC was different from that of IVT+ among patients without colorectal neoplasies. The occurrence of IVT+ in skin fibroblasts among patients with CRC was associated with the occurrence of occult mandibular osteomas in the same patients. The occurrence of dermal fibroblast IVT+ conveyed any prognostic significance in CRC patients. RESULTS AND CONCLUSIONS: 1) A significant correlation was found between both the occurrence of skin fibroblast IVT+ and adenoma DNA aneuploidy in relation to the degree of dysplasia and histological type of adenomas. This signifies, that inborn and acquired genetic instability is correlated to the adenoma-carcinoma sequence. IVT+ was found to be correlated to the progression of adenomas to carcinomas and not to the development of adenomas. 2) A direct correlation between IVT+ and adenoma DNA aneuploidy could not be demonstrated. However, among those patients with diploid adenomas and IVT+ 57% showed villous adenomas. Among those patients with diploid adenomas and IVT- only 14% had villous adenomas. This further substantiates the correlation of IVT+ to the adenoma-carcinoma sequence. IVT+ was highly associated with DNA aneuploidy in carcinomas, and since IVT+ was found to be significantly associated with CRC's with a DNA index > or = 1.5, it is suggested that IVT+ mainly is correlated to the early steps in tumor progression. 3) IVT+ was found in 34% of sporadic CRC, and tumor DNA aneuploidy was demonstrated in 73%...

Adenoma↗

Classical oncogenes and tumor suppressor genes: a comparative genomics perspective.

We have curated a reference set of cancer- related genes and reanalyzed their sequences in the light of molecular information and resources that have become available since they were first cloned. Homology studies were carried out for human oncogenes and tumor suppressors, compared with the complete proteome of the nematode, Caenorhabditis elegans, and partial proteomes of mouse and rat and the fruit fly, Drosophila melanogaster. Our results demonstrate that simple, semi-automated bioinformatics approaches to identifying putative functionally equivalent gene products in different organisms may often be misleading. An electronic supplement to this article provides an integrated view of our comparative genomics analysis as well as mapping data, physical cDNA resources and links to published literature and reviews, thus creating a "window" into the genomes of humans and other organisms for cancer biology.

Animals↗

Characterization of the Ac/Ds behaviour in transgenic tomato plants using plasmid rescue.

We describe the use of plasmid rescue to facilitate studies on the behaviour of Ds and Ac elements in transgenic tomato plants. The rescue of Ds elements relies on the presence of a plasmid origin of replication and a marker gene selective in Escherichia coli within the element. The position within the genome of modified Ds elements, rescued both before and after transposition, is assigned to the RFLP map of tomato. Alternatively to the rescue of Ds elements equipped with plasmid sequences, Ac elements are rescued by virtue of plasmid sequences flanking the element. In this way, the consequences of the presence of an (active) Ac element on the DNA structure at the original site can be studied in detail. Analysis of a library of Ac elements, rescued from the genome of a primary transformant, shows that Ac elements are, infrequently, involved in the formation of deletions. In one case the deletion refers to a 174 bp genomic DNA sequence immediately flanking Ac. In another case, a 1878 bp internal Ac sequence is deleted.

Base Sequence↗

Persistent infections and immunity in cystic fibrosis.

Cystic fibrosis (CF) is the most common autosomal recessive lethal disease in the Caucasian population. Chronic respiratory infections with Pseudomonas aeruginosa, neutrophil-dominated airway inflammation and progressive lung damage are the major causes of morbidity and mortality in CF. Two persistent infection phenotypes expressed by this bacterium are biofilm and mucoidy. Biofilm, also called the microcolony mode of growth is the surface-associated adherent bacterial community, while mucoidy refers to a phenotype conducive to copious amounts of mucoid exopolysaccharide (MEP)/alginate that provides a matrix for mature biofilms conferring resistance to host defenses and antibiotics. Recent completion of the whole genomic sequence of the standard reference strain P. aeruginosa PAO1 has led to discoveries that many clinical isolates of this species possess unique genomic sequences (genomic islands) due to horizontal gene transfer. We propose this type of genetic exchange may play an important role in causing intrinsic genomic diversity of this organism. Therefore, the diversity, as revealed through profiles of restriction fragment length polymorphism (RFLP), may be linked to an array of novel and unexplored pathogenic mechanisms in P. aeruginosa. CF mouse models, while displaying many clinical similarities to human CF, have yet to demonstrate a chronic pulmonary disease phenotype. This review is intended to provide an overview of P. aeruginosa persistent infection phenotypes (biofilm and mucoidy) and an aerosol infection mouse model for CF. Genomic diversity of P. aeruginosa and its implications in the pathogenesis in CF will also be discussed.

Animals↗

Compilation and characterization of a novel WNK family of protein kinases in Arabiodpsis thaliana with reference to circadian rhythms.

The complete genome sequence of Arabidopsis thaliana revealed that this higher plant has a tremendous number of protein kinases. We recently isolated a novel type of protein kinase, named AtWNK1, which shows an in vitro ability to phosphorylate the APRR3 member of the APRR1/TOC1 quintet that has been implicated in a mechanism underlying circadian rhythms in Arabidopsis. We here address two issues, one general and one specific, as to this novel protein kinase. We first asked the general question of how many WNK family members are present in this higher plant, then whether or not other members are also relevant to circadian rhythms. The results of our analyses showed that Arabidopsis has at least 9 members of the WNK1 family of protein kinases (designated here as WNK1 to WNK9), the structural design of which is clearly distinct from those of other known protein kinases, such as receptor-like kinases and mitogen-activated protein kinases. They were examined with special reference to the circadian-related APRR1/TOC1 quintet. It was found that not only the transcription of the WNK1 gene, but also those of three other members (WNK2, WNK4, and WNK6) are under the control of circadian rhythms. These results suggested that certain members of the WNK family of protein kinases might play roles in a mechanism that generates circadian rhythms in Arabidopsis.

Amino Acid Sequence↗

Characterization and comparison of the DNAs of the three closely related bacteriophages gd, ge and gf with the genome DNA of the hydrogen-oxidizing host strain Pseudomonas pseudoflava GA3.

The double-stranded (ds)DNAs of the three closely related temperate Pseudomonas pseudoflava bacteriophages gd, ge and gf were studied biochemically and biophysically. The GC content of the DNA was 67.4 +/- 0.5% and differed only slightly from that of the host P. pseudoflava. By electron microscopic length measurements a mol. wt. of 26.1 X 10(6) to 26.7 X 10(6) was calculated for the three bacteriophage DNAs. Homogeneity of the bacteriophage DNAs was further demonstrated by specific cleavage with restriction endonucleases WcoRI and HindIII. It was concluded that the three homo-immune bacteriophages are identical. The genome size of the host P. pseudoflava GA3 was 3.7 X 10(9) as calculated from optical renaturation rate measurement with Xanthomonas pelargonii reference DNA. The bacteriophage gd genome thus amounts to 0.7% of the chromosome of this bacterial host.

Bacteriophages↗

Comparative genomic hybridization-array analysis enhances the detection of aneuploidies and submicroscopic imbalances in spontaneous miscarriages.

Miscarriage is a condition that affects 10%-15% of all clinically recognized pregnancies, most of which occur in the first trimester. Approximately 50% of first-trimester miscarriages result from fetal chromosome abnormalities. Currently, G-banded chromosome analysis is used to determine if large-scale genetic imbalances are the cause of these pregnancy losses. This technique relies on the culture of cells derived from the fetus, a technique that has many limitations, including a high rate of culture failure, maternal overgrowth of fetal cells, and poor chromosome morphology. Comparative genomic hybridization (CGH)-array analysis is a powerful new molecular cytogenetic technique that allows genomewide analysis of DNA copy number. By hybridizing patient DNA and normal reference DNA to arrays of genomic clones, unbalanced gains or losses of genetic material across the genome can be detected. In this study, 41 product-of-conception (POC) samples, which were previously analyzed by G-banding, were tested using CGH arrays to determine not only if the array could identify all reported abnormalities, but also whether any previously undetected genomic imbalances would be discovered. The array methodology detected all abnormalities as reported by G-banding analysis and revealed new abnormalities in 4/41 (9.8%) cases. Of those, one trisomy 21 POC was also mosaic for trisomy 20, one had a duplication of the 10q telomere region, one had an interstitial deletion of chromosome 9p, and the fourth had an interstitial duplication of the Prader-Willi/Angelman syndrome region on chromosome 15q, which, if maternally inherited, has been implicated in autism. This retrospective study demonstrates that the DNA-based CGH-array technology overcomes many of the limitations of routine cytogenetic analysis of POC samples while enhancing the detection of fetal chromosome aberrations.

Abortion, Spontaneous↗

TFBScluster web server for the identification of mammalian composite regulatory elements.

Identification of transcriptional regulatory elements represents a critical step in our ability to reconstruct transcriptional regulatory networks from gene expression profiling datasets. To facilitate computational identification of candidate gene regulatory elements from whole genome sequences, we have developed the TFBScluster web server that integrates several tools for the genome-wide identification and subsequent characterization of transcription factor binding site clusters that are conserved in multiple mammalian species. Either the human or mouse genomes can be used as the reference sequence with direct links from the search results to the ENSEMBL and UCSC genome browsers. Moreover, TFBScluster provides seamless integration of transcription factor binding site searches with genome annotation and gene expression profiling data, to allow prioritising computational predictions for subsequent experimental validation. TFBScluster is publicly available at http://hscl.cimr.cam.ac.uk/TFBScluster_genome_portal.html.

Animals↗

Transmissible enteritis of turkeys: experimental inoculation studies with tissue-culture-adapted turkey and bovine coronaviruses.

Four Quebec isolates of turkey enteric coronaviruses (TCVs) and three isolates of bovine enteric coronaviruses (BCVs) were serially propagated in HRT-18 and compared for their pathogenicity in turkey embryos and turkey poults. By immunoelectron microscopy, hemagglutination-inhibition, and Western immunoblotting assays, tissue-culture-adapted Quebec TCV isolates were found to be closely related to the reference Minnesota strain of TCV and the Mebus strain of BCV. Genomic relationships between TCV isolates and the reference BCV strain were confirmed by hybridization assays with BCV-specific radiolabeled recombinant plasmids containing sequences of the N and M genes. Only TCV isolates could be propagated by inoculation in the amniotic cavity of chicken and turkey embryonating eggs, and induced clinical disease in turkey poults. Nevertheless, coronavirus particles or antigens were detected by electron microscopy or indirect enzyme-linked immunosorbent assay in the clarified intestinal contents of BCV-infected poults up to day 14 PI, and genomic viral RNA was detected by slot-blot hybridization using BCV cDNA probes.

Animals↗

MetagenomicKG: a knowledge graph for metagenomic applications.

MOTIVATION: The sheer volume and variety of genomic content within microbial communities makes metagenomics a field rich in biomedical knowledge. To traverse these complex communities and their vast unknowns, metagenomic studies often depend on distinct reference databases, such as the Genome Taxonomy Database (GTDB), the Kyoto Encyclopedia of Genes and Genomes (KEGG), and the Bacterial and Viral Bioinformatics Resource Center (BV-BRC), for various analytical purposes. These databases are crucial for the genetic and functional annotation of microbial communities. Nevertheless, the inconsistent nomenclature or identifiers of these databases present challenges for effective integration, representation, and utilization. Knowledge graphs (KGs) offer an appropriate solution by organizing biological entities from different databases to standardized identifiers, allowing their interrelations to be captured into a cohesive network regardless of the naming conventions used in each source. The graph structure not only facilitates the unveiling of hidden patterns but also enriches our biological understanding with deeper insights. Despite KGs having shown potential in various biomedical fields, their application in metagenomics remains underexplored. RESULTS: We present MetagenomicKG, a novel knowledge graph specifically tailored for metagenomic analysis. MetagenomicKG integrates taxonomic, functional, and pathogenesis-related information on the human microbiome sourced from various databases, and further connects these with existing biomedical KGs to expand the biological network. Through various case studies involving the human microbiome, we demonstrate its utility in enabling hypothesis generation regarding the relationships between microbes and diseases, generating sample-specific graph embeddings, and providing robust pathogen prediction. CODE AVAILABILITY: The source code and technical details for constructing the MetagenomicKG and reproducing all analyses are available on GitHub at https://github.com/KoslickiLab/MetagenomicKG. The data used in this manuscript, including the pre-built files and use case input data, are archived on Zenodo with DOI: 10.5281/zenodo.17546861.

Metagenomics↗