Search PubMed⌕ Search

Biomedical subjects

Ikuo Uchiyama

Publications and source records attributed to Ikuo Uchiyama.

14 recordsLinked to original sources

Microbiome Datahub: an open-access platform integrating environmental metadata, taxonomy, and functional annotation for comprehensive metagenome-assembled genome datasets.

BACKGROUND: Metagenome-assembled genomes (MAGs) provide crucial insights into the genomic diversity of uncultured microbes. However, MAG datasets deposited in public repositories such as INSDC are often difficult to reuse due to heterogeneous quality, inconsistent taxonomic and functional annotations, and insufficiently curated environmental metadata. While secondary MAG databases such as MGnify, IMG/M, and SPIRE provide standardized resources, they reconstruct MAGs de novo from public metagenomic reads and therefore do not represent the original MAGs reported in publications. RESULTS: To address this gap, we developed Microbiome Datahub, an open-access platform that systematically aggregates and re-annotates original MAGs from INSDC. We collected 214,427 MAGs, predicted genes by DFAST, performed quality assessment with CheckM, standardized taxonomic assignments with GTDB-Tk, inferred 27 phenotypic traits using Bac2Feature, assigned proteins to MBGD ortholog clusters and KEGG Orthology IDs using PZLAST, and annotated environmental metadata with the Metagenome and Microbes Environmental Ontology. Across these MAGs, the average completeness was 80.5% and contamination 1.8%; notably, the most frequent values were&#x2009;>95% completeness and&#x2009;<1% contamination, indicating that the majority of MAGs are of high quality. Comparative analyses showed that Microbiome Datahub provides phylogenetically and environmentally diverse MAGs: while the majority originated from vertebrate gut environments, a substantial number were also recovered from other habitats such as groundwater, including nearly 10,000 MAGs from the Patescibacteria. Inference of 27 phenotypic traits, including optimum growth temperature, further revealed ecological differentiation across phyla. Protein clustering revealed 56 million identity 40% clusters, with the majority unique compared with MGnify and GlobDB, and&#x2009;~19% of proteins unassigned to MBGD ortholog clusters, underscoring their novelty. CONCLUSIONS: Microbiome Datahub integrates MAG genome sequences, gene and protein predictions, quality metrics, environmental and taxonomic annotations, ortholog cluster assignments, and phenotype predictions, all accessible via a web interface, API, and bulk downloads. By combining original MAGs with curated metadata and functional annotations, Microbiome Datahub constitutes a comprehensive and reusable resource that will accelerate microbiome and microbial genomics research. Video Abstract.

Metagenome↗

MBGD: a platform for microbial comparative genomics based on the automated construction of orthologous groups.

The microbial genome database for comparative analysis (MBGD) is a comprehensive platform for microbial comparative genomics. The central function of MBGD is to create orthologous groups among multiple genomes from precomputed all-against-all similarity relationships using the DomClust algorithm. The database now contains >300 published genomes and the number continues to grow. For researchers who are interested in ongoing genome projects, we have now started a new service called 'My MBGD,' which allows users to add their own genome sequences to MBGD for the purpose of identifying orthologs among both the new and the existing genomes. Furthermore, in order to make available the rapidly accumulating information on closely related genome sequences, we enhanced the interface for pairwise genome comparisons using the CGAT interface, which allows users to see nucleotide sequence alignments of non-coding as well as coding regions. MBGD is available at http://mbgd.genome.ad.jp/.

Algorithms↗

CGAT: a comparative genome analysis tool for visualizing alignments in the analysis of complex evolutionary changes between closely related genomes.

BACKGROUND: The recent accumulation of closely related genomic sequences provides a valuable resource for the elucidation of the evolutionary histories of various organisms. However, although numerous alignment calculation and visualization tools have been developed to date, the analysis of complex genomic changes, such as large insertions, deletions, inversions, translocations and duplications, still presents certain difficulties. RESULTS: We have developed a comparative genome analysis tool, named CGAT, which allows detailed comparisons of closely related bacteria-sized genomes mainly through visualizing middle-to-large-scale changes to infer underlying mechanisms. CGAT displays precomputed pairwise genome alignments on both dotplot and alignment viewers with scrolling and zooming functions, and allows users to move along the pre-identified orthologous alignments. Users can place several types of information on this alignment, such as the presence of tandem repeats or interspersed repetitive sequences and changes in G+C contents or codon usage bias, thereby facilitating the interpretation of the observed genomic changes. In addition to displaying precomputed alignments, the viewer can dynamically calculate the alignments between specified regions; this feature is especially useful for examining the alignment boundaries, as these boundaries are often obscure and can vary between programs. Besides the alignment browser functionalities, CGAT also contains an alignment data construction module, which contains various procedures that are commonly used for pre- and post-processing for large-scale alignment calculation, such as the split-and-merge protocol for calculating long alignments, chaining adjacent alignments, and ortholog identification. Indeed, CGAT provides a general framework for the calculation of genome-scale alignments using various existing programs as alignment engines, which allows users to compare the outputs of different alignment programs. Earlier versions of this program have been used successfully in our research to infer the evolutionary history of apparently complex genome changes between closely related eubacteria and archaea. CONCLUSION: CGAT is a practical tool for analyzing complex genomic changes between closely related genomes using existing alignment programs and other sequence analysis tools combined with extensive manual inspection.

Algorithms↗

How genomes rearrange: genome comparison within bacteria Neisseria suggests roles for mobile elements in formation of complex genome polymorphisms.

Comparison of closely related genome sequences can provide a clue as to how macroscopic genome polymorphisms were formed through various events of recombination. However, this approach has been limited to relatively simple polymorphisms such as insertion, deletion and inversion. In the present study, we tried to extend this approach to more complex genome polymorphisms that were observed when four genome sequences of bacterial genus Neisseria were compared. The first polymorphism was an apparent translocation (ab-cd to cd-ba; a region 'ab' was translocated). The second one was a re-ordering of adjacent regions (ab-cd-ef-gh to ef-cd-ab-gh; ab, cd and ef were in reverse order). The third one was a translocation of two adjacent regions with permutation of their order (ab-cd to cd-ab elsewhere in the genome). The fourth one was a genome-wide inversion associated with a genome-specific insertion into the joints (-ab-cd- to -y-ba-x-cd-). We were able to explain their formation by only a few steps of plausible events of recombination that involved linked IS copies and prophages. Our approach would help to reconstruct a history of apparently complex genome polymorphisms in any forms of organisms and to understand genome rearrangements in the natural environments in non-model organisms.

Base Sequence↗

Evolution of paralogous genes: Reconstruction of genome rearrangements through comparison of multiple genomes within Staphylococcus aureus.

Analysis of evolution of paralogous genes in a genome is central to our understanding of genome evolution. Comparison of closely related bacterial genomes, which has provided clues as to how genome sequences evolve under natural conditions, would help in such an analysis. With species Staphylococcus aureus, whole-genome sequences have been decoded for seven strains. We compared their DNA sequences to detect large genome polymorphisms and to deduce mechanisms of genome rearrangements that have formed each of them. We first compared strains N315 and Mu50, which make one of the most closely related strain pairs, at the single-nucleotide resolution to catalogue all the middle-sized (more than 10 bp) to large genome polymorphisms such as indels and substitutions. These polymorphisms include two paralogous gene sets, one in a tandem paralogue gene cluster for toxins in a genomic island and the other in a ribosomal RNA operon. We also focused on two other tandem paralogue gene clusters and type I restriction-modification (RM) genes on the genomic islands. Then we reconstructed rearrangement events responsible for these polymorphisms, in the paralogous genes and the others, with reference to the other five genomes. For the tandem paralogue gene clusters, we were able to infer sequences for homologous recombination generating the change in the repeat number. These sequences were conserved among the repeated paralogous units likely because of their functional importance. The sequence specificity (S) subunit of type I RM systems showed recombination, likely at the homology of a conserved region, between the two variable regions for sequence specificity. We also noticed novel alleles in the ribosomal RNA operons and suggested a role for illegitimate recombination in their formation. These results revealed importance of recombination involving long conserved sequence in the evolution of paralogous genes in the genome.

Amino Acid Sequence↗

Genome comparison in silico in Neisseria suggests integration of filamentous bacteriophages by their own transposase.

We have identified filamentous prophages, Nf (Neisserial filamentous phages), during an in silico genome comparison in Neisseria. Comparison of three genomes of Neisseria meningitidis and one of Neisseria gonorrhoeae revealed four subtypes of Nf. Eleven intact copies are located at different loci in the four genomes. Each intact copy of Nf is flanked by duplication of 5'-CT and, at its right end, carries a transposase homologue (pivNM/irg) of RNaseH/Retroviral integrase superfamily. The phylogeny of these putative transposases and that of phage-related proteins on Nfs are congruent. Following circularization of Nfs, a promoter-like sequence forms. The sequence at the junction of these predicted circular forms (5'-atCTtatat) was found in a related plasmid (pMU1) at a corresponding locus. Several structural variants of Nfs--partially inverted, internally deleted and truncated--were also identified. The partial inversion seems to be a product of site-specific recombination between two 5'-CTtat sequences that are in inverse orientation, one at its end and the other upstream of pivNM/irg. Formation of internally deleted variants probably proceeded through replicative transposition that also involved two 5'-CTtat sequences. We concluded that the PivNM/Irg transposase on Nfs integrated their circular forms into the chromosomal 5'-CT-containing sequences and probably mediated the above rearrangements.

Base Sequence↗

Hierarchical clustering algorithm for comprehensive orthologous-domain classification in multiple genomes.

Ortholog identification is a crucial first step in comparative genomics. Here, we present a rapid method of ortholog grouping which is effective enough to allow the comparison of many genomes simultaneously. The method takes as input all-against-all similarity data and classifies genes based on the traditional hierarchical clustering algorithm UPGMA. In the course of clustering, the method detects domain fusion or fission events, and splits clusters into domains if required. The subsequent procedure splits the resulting trees such that intra-species paralogous genes are divided into different groups so as to create plausible orthologous groups. As a result, the procedure can split genes into the domains minimally required for ortholog grouping. The procedure, named DomClust, was tested using the COG database as a reference. When comparing several clustering algorithms combined with the conventional bidirectional best-hit (BBH) criterion, we found that our method generally showed better agreement with the COG classification. By comparing the clustering results generated from datasets of different releases, we also found that our method showed relatively good stability in comparison to the BBH-based methods.

Algorithms↗

Discovery of a novel restriction endonuclease by genome comparison and application of a wheat-germ-based cell-free translation assay: PabI (5'-GTA/C) from the hyperthermophilic archaeon Pyrococcus abyssi.

To search for restriction endonucleases, we used a novel plant-based cell-free translation procedure that bypasses the toxicity of these enzymes. To identify candidate genes, the related genomes of the hyperthermophilic archaea Pyrococcus abyssi and Pyrococcus horikoshii were compared. In line with the selfish mobile gene hypothesis for restriction-modification systems, apparent genome rearrangement around putative restriction genes served as a selecting criterion. Several candidate restriction genes were identified and then amplified in such a way that they were removed from their own translation signal. During their cloning into a plasmid, the genes became connected with a plant translation signal. After in vitro transcription by T7 RNA polymerase, the mRNAs were separated from the template DNA and translated in a wheat-germ-based cell-free protein synthesis system. The resulting solution could be directly assayed for restriction activity. We identified two deoxyribonucleases. The novel enzyme was denoted as PabI, purified and found to recognize 5'-GTAC and leave a 3'-TA overhang (5'-GTA/C), a novel restriction enzyme-generated terminus. PabI is active up to 90 degrees C and optimally active at a pH of around 6 and in NaCl concentrations ranging from 100 to 200 mM. We predict that it has a novel 3D structure.

Base Sequence↗

Analysis of expressed sequence tags of the water flea Daphnia magna.

To study gene expression in the water flea Daphnia magna we constructed a cDNA library and characterized the expressed sequence tags (ESTs) of 7210 clones. The EST sequences clustered into 2958 nonredundant groups. BLAST analyses of both protein and DNA databases showed that 1218 (41%) of the unique sequences shared significant similarities to known nucleotide or amino acid sequences, whereas the remaining 1740 (59%) showed no significant similarities to other genes. Clustering analysis revealed particularly high expression of genes related to ATP synthesis, structural proteins, and proteases. The cDNA clones and EST sequence information should be useful for future functional analysis of daphnid biology and investigation of the links between ecology and genomics.

Animals↗

Thermoadaptation trait revealed by the genome sequence of thermophilic Geobacillus kaustophilus.

We present herein the first complete genome sequence of a thermophilic Bacillus-related species, Geobacillus kaustophilus HTA426, which is composed of a 3.54 Mb chromosome and a 47.9 kb plasmid, along with a comparative analysis with five other mesophilic bacillar genomes. Upon orthologous grouping of the six bacillar sequenced genomes, it was found that 1257 common orthologous groups composed of 1308 genes (37%) are shared by all the bacilli, whereas 839 genes (24%) in the G.kaustophilus genome were found to be unique to that species. We were able to find the first prokaryotic sperm protamine P1 homolog, polyamine synthase, polyamine ABC transporter and RNA methylase in the 839 unique genes; these may contribute to thermophily by stabilizing the nucleic acids. Contrasting results were obtained from the principal component analysis (PCA) of the amino acid composition and synonymous codon usage for highlighting the thermophilic signature of the G.kaustophilus genome. Only in the PCA of the amino acid composition were the Bacillus-related species located near, but were distinguishable from, the borderline distinguishing thermophiles from mesophiles on the second principal axis. Further analysis revealed some asymmetric amino acid substitutions between the thermophiles and the mesophiles, which are possibly associated with the thermoadaptation of the organism.

Adaptation, Physiological↗

Comparative genomics of Physcomitrella patens gametophytic transcriptome and Arabidopsis thaliana: implication for land plant evolution.

The mosses and flowering plants diverged >400 million years ago. The mosses have haploid-dominant life cycles, whereas the flowering plants are diploid-dominant. The common ancestors of land plants have been inferred to be haploid-dominant, suggesting that genes used in the diploid body of flowering plants were recruited from the genes used in the haploid body of the ancestors during the evolution of land plants. To assess this evolutionary hypothesis, we constructed an EST library of the moss Physcomitrella patens, and compared the moss transcriptome to the genome of Arabidopsis thaliana. We constructed full-length enriched cDNA libraries from auxin-treated, cytokinin-treated, and untreated gametophytes of P. patens, and sequenced both ends of >40,000 clones. These data, together with the mRNA sequences in the public databases, were assembled into 15,883 putative transcripts. Sequence comparisons of A. thaliana and P. patens showed that at least 66% of the A. thaliana genes had homologues in P. patens. Comparison of the P. patens putative transcripts with all known proteins, revealed 9,907 putative transcripts with high levels of similarity to vascular plant genes, and 850 putative transcripts with high levels of similarity to other organisms. The haploid transcriptome of P. patens appears to be quite similar to the A. thaliana genome, supporting the evolutionary hypothesis. Our study also revealed that a number of genes are moss specific and were lost in the flowering plant lineage.

Arabidopsis↗

MBGD: microbial genome database for comparative analysis.

MBGD is a workbench system for comparative analysis of completely sequenced microbial genomes. The central function of MBGD is to create an orthologous gene classification table using precomputed all-against-all similarity relationships among genes in multiple genomes. In MBGD, an automated classification algorithm has been implemented so that users can create their own classification table by specifying a set of organisms and parameters. This feature is especially useful when the user's interest is focused on some taxonomically related organisms. The created classification table is stored into the database and can be explored combining with the data of individual genomes as well as similarity relationships among genomes. Using these data, users can carry out comparative analyses from various points of view, such as phylogenetic pattern analysis, gene order comparison and detailed gene structure comparison. MBGD is accessible at http://mbgd.genome.ad.jp/.

Algorithms↗

Genome sequence of Oceanobacillus iheyensis isolated from the Iheya Ridge and its unexpected adaptive capabilities to extreme environments.

Oceanobacillus iheyensis HTE831 is an alkaliphilic and extremely halotolerant Bacillus-related species isolated from deep-sea sediment. We present here the complete genome sequence of HTE831 along with analyses of genes required for adaptation to highly alkaline and saline environments. The genome consists of 3.6 Mb, encoding many proteins potentially associated with roles in regulation of intracellular osmotic pressure and pH homeostasis. The candidate genes involved in alkaliphily were determined based on comparative analysis with three Bacillus species and two other Gram-positive species. Comparison with the genomes of other major Gram-positive bacterial species suggests that the backbone of the genus Bacillus is composed of approximately 350 genes. This second genome sequence of an alkaliphilic Bacillus-related species will be useful in understanding life in highly alkaline environments and microbial diversity within the ubiquitous bacilli.

Bacillus↗