Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Genome mining for new enediyne antibiotics.

Enediyne antibiotics epitomize nature's chemical creativity. They contain intricate molecular architectures that are coupled with potent biological activities involving double-stranded DNA scission. The recent explosion in microbial genome sequences has revealed a large reservoir of novel enediynes. However, while hundreds of enediyne biosynthetic gene clusters (BGCs) can be detected, less than two dozen natural products have been characterized to date as many clusters remain silent or sparingly expressed under standard laboratory growth conditions. This review focuses on four distinct strategies, which have recently enabled discoveries of novel enediynes: phenotypic screening from rare sources, biosynthetic manipulation, genomic signature-based PCR screening, and DNA-cleavage assays coupled with activation of silent BGCs via high-throughput elicitor screening. With an abundance of enediyne BGCs and emerging approaches for accessing them, new enediyne natural products and further insights into their biogenesis are imminent.

Enediynes↗

Discovery of a thermostable Baeyer-Villiger monooxygenase by genome mining.

Baeyer-Villiger monooxygenases represent useful biocatalytic tools, as they can catalyze reactions which are difficult to achieve using chemical means. However, only a limited number of these atypical monooxygenases are available in recombinant form. Using a recently described protein sequence motif, a putative Baeyer-Villiger monooxygenase (BVMO) was identified in the genome of the thermophilic actinomycete Thermobifida fusca. Heterologous expression of the respective protein in Escherichia coli and subsequent enzyme characterization showed that it indeed represents a BVMO. The NADPH-dependent and FAD-containing monooxygenase is active with a wide range of aromatic ketones, while aliphatic substrates are also converted. The best substrate discovered so far is phenylacetone (k(cat) = 1.9 s(-1), K(M) = 59 microM). The enzyme exhibits moderate enantioselectivity with alpha-methylphenylacetone (enantiomeric ratio of 7). In addition to Baeyer-Villiger reactions, the enzyme is able to perform sulfur oxidations. Different from all known BVMOs, this newly identified biocatalyst is relatively thermostable, displaying an activity half-life of 1 day at 52 degrees C. This study demonstrates that, using effective annotation tools, genomes can efficiently be exploited as a source of novel BVMOs.

Actinomycetales↗

Genomic mining type III secretion system effectors in Pseudomonas syringae yields new picks for all TTSS prospectors.

Many bacterial pathogens of plants and animals use a type III secretion system (TTSS) to deliver virulence effector proteins into host cells. Because effectors are heterogeneous in sequence and function, there has not been a systematic way to identify the genes encoding them in pathogen genomes, and our current inventories are probably incomplete. A pre-closure draft sequence of Pseudomonas syringae pv. tomato DC3000, a pathogen of tomato and Arabidopsis, has recently supported five complementary studies which, collectively, identify 36 TTSS-secreted proteins and many more candidate effectors in this strain. These studies demonstrate the advantages of combining experimental and computational approaches, and they yield new insights into TTSS effectors and virulence regulation in P. syringae, potential effector targeting signals in all TTSS-dependent pathogens, and strategies for finding TTSS effectors in other bacteria that have sequenced genomes.

Arabidopsis↗

Genome mining of alkaliphilic cyanobacterial consortia: identification of biosynthetic gene clusters in Sodalinema and associated heterotrophs.

Alkaline soda lakes are high-pH environments that host specialized microbial communities with potential for biotechnology and natural product discovery. We characterized three Sodalinema-dominated cyanobacterial consortia enriched from Canadian soda lakes over 510 days. Using hybrid metagenomic sequencing and metatranscriptomics across pH, alkalinity, and temperature gradients, we reconstructed high-quality metagenome-assembled genomes and assessed functional activity. All consortia converged toward cyanobacteria dominance and exhibited temperature optima between 21°C and 30°C. Phylogenetic analysis placed Sodalinema genomes within a distinct clade affiliated with Candidatus Sodalinema alkaliphilum. Genomic analysis indicated complete biosynthetic pathways for vitamin B5, vitamin B7, and the molybdenum cofactor, but incomplete pathways for vitamins B1, B9, and B12, consistent with patterns observed in Sodalinema yuhuli. Metatranscriptomic profiles showed increased expression of genes involved in phycocyanin and carotenoid biosynthesis at pH 10.2 relative to pH 8.5. Biosynthetic gene cluster analysis revealed that most secondary metabolic potential resided in heterotrophic community members. Roseinatronobacter encoded pathways for N-acyl homoserine lactones, osmoprotectants, betalactones, and prodigiosin, while Alkalimonas, Wenzhouxiangella, and members of the Kiloniellales encoded clusters for lanthipeptides, cyclodipeptides, hydrogen cyanide, and pyrroloquinoline quinone. These findings indicate functional partitioning within the consortia and highlight the contribution of heterotrophs to secondary metabolism.IMPORTANCEAlkaline soda lakes contain microbial communities adapted to high pH that remain underexplored for biotechnology. This study focuses on Sodalinema, a filamentous cyanobacterium that dominates enriched consortia from Canadian soda lakes, and its associated heterotrophic partners. We show that while Sodalinema drives primary productivity, heterotrophic bacteria encode most of the pathways for antimicrobial and signaling compounds. These interactions may support community stability and defense against competing microorganisms. By linking genomic potential with gene expression, this work identifies alkaline cyanobacterial consortia as a source of bioactive compounds and provides a framework for exploring extremophilic microbial communities for natural product discovery.

Sodalinema↗

BAGEL: a web-based bacteriocin genome mining tool.

A common problem in the annotation of open reading frames (ORFs) is the identification of genes that are functionally similar but have limited or no sequence homology. This is particularly the case for bacteriocins, a very diverse group of antimicrobial peptides produced by bacteria and usually encoded by small, poorly conserved ORFs. ORFs surrounding bacteriocin genes are often biosynthetic genes. This information can be used to locate putative structural bacteriocin genes. Here, we describe BAGEL, a web server that identifies putative bacteriocin ORFs in a DNA sequence using novel, knowledge-based bacteriocin databases and motif databases. Many bacteriocins are encoded by small genes that are often omitted in the annotation process of bacterial genomes. Thus, we have implemented ORF detection using a number of published ORF prediction tools. In addition, BAGEL takes into account the genomic context, i.e. for each potential bacteriocin-encoding ORF, the sequence of the surrounding region on the genome is analyzed for genes that might encode proteins involved in biosynthesis, transport, regulation and/or immunity. These innovations make BAGEL unique in its ability to detect putative bacteriocin gene clusters in (new) bacterial genomes. BAGEL is freely accessible at: http://bioinformatics.biol.rug.nl/websoftware/bagel.

Amino Acid Motifs↗

Genomic mining of new genes and pathways in innate and adaptive immunity.

Plant disease resistant (R) genes constitute a large family that mediates host response to bacteria, viruses and fungi. Large mammalian proteins containing a nucleotide binding domain (NBD) and C-terminal leucine-rich repeats (LRRs) are similar in structure to the TLR/NBD/LRR subfamily of R proteins and have been suggested as a link between innate and acquired immunity. Because of our long-term interest in one of these, the class II transactivator (CIITA), and recent reports linking mutations in two new NBD/LRR proteins (Nod2/CARD 15 and CIAS 1/cryopyrin) to various autoimmune and inflammatory disorders, we have performed a comprehensive search of the human genome and found a multigene family which we termed the CATERPILLAR (CARD, Transcription Enhancer, R[purine]-binding, Pyrin, Lots of Leucine Regions)family. The N-termini of these genes are varied although the majority have a pyrin domain and few have a CARD domain. The genomic organization of these genes demonstrates a high degree of conservation with the NBD encoded as a single large exon and the LRRs encoded in two basic arrangements. Detailed analysis and new functional data regarding a number of the CATERPILLAR proteins will be described, including CIITA, cryopyrin and Monarch 1.

Adaptation, Physiological↗

Genome mining in Streptomyces coelicolor: molecular cloning and characterization of a new sesquiterpene synthase.

The terpene synthase encoded by the SCO5222 (SC7E4.19) gene of Streptomyces coelicolor was cloned by PCR and expressed in Escherichia coli as an N-terminal-His6-tag protein. Incubation of the recombinant protein, SCO5222p, with farnesyl diphosphate (1, FPP) in the presence of Mg(II) gave a new sesquiterpene, (+)-epi-isozizaene (2), whose structure and stereochemistry were determined by a combination of 1H, 13C, COSY, HMQC, HMBC, and NOESY NMR. The steady-state kinetic parameters were kcat 0.049 +/- 0.001 s-1 and a Km (FPP) of 147 +/- 14 nM. Individual incubations of recombinant epi-isozizaene synthase with [1,1-2H2]FPP (1a), (1R)-[1-2H]-FPP (1b), and (1S)-[1-2H]-FPP (1c) and NMR analysis of the resulting deuterated epi-isozizaenes supported an isomerization-cyclization-rearrangement mechanism involving the intermediacy of (3R)-nerolidyl diphosphate (3).

Alkyl and Aryl Transferases↗

Type III polyketide synthase beta-ketoacyl-ACP starter unit and ethylmalonyl-CoA extender unit selectivity discovered by Streptomyces coelicolor genome mining.

Polyketide synthases (PKSs) are involved in the biosynthesis of many important natural products. In bacteria, type III PKSs typically catalyze iterative decarboxylation and condensation reactions of malonyl-CoA building blocks in the biosynthesis of polyhydroxyaromatic products. Here it is shown that Gcs, a type III PKS encoded by the sco7221 ORF of the bacterium Streptomyces coelicolor, is required for biosynthesis of the germicidin family of 3,6-dialkyl-4-hydroxypyran-2-one natural products. Evidence consistent with Gcs-catalyzed elongation of specific beta-ketoacyl-ACP products of the fatty acid synthase FabH with ethyl- or methylmalonyl-CoA in the biosynthesis of germicidins is presented. Selectivity for beta-ketoacyl-ACP starter units and ethylmalonyl-CoA as an extender unit is unprecedented for type III PKSs, suggesting these enzymes may be capable of utilizing a far wider range of starter and extender units for natural product assembly than believed until now.

3-Oxoacyl-(Acyl-Carrier-Protein) Synthase↗

Helminth vaccines: from mining genomic information for vaccine targets to systems used for protein expression.

The control of helminth diseases of people and livestock continues to rely on the widespread use of anti-helminthic drugs. However, concerns with the appearance of drug resistant parasites and the presence of pesticide residues in food and the environment, has given further incentive to the goal of discovering molecular vaccines against these pathogens. The exponential rate at which gene and protein sequence information is accruing for many helminth parasites requires new methods for the assimilation and analysis of the data and for the identification of molecules capable of inducing immunological protection. Some promising vaccine candidates have been discovered, in particular cathepsin L proteases from Fasciola hepatica, aminopeptidases from Haemonchus contortus, and aspartic proteases from schistosomes and hookworms, all of which are secreted into the host tissues or into the parasite intestine where they play important roles in host-parasite interactions. Since secreted proteins, in general, are exposed to the immune system of the host they represent obvious candidates at which vaccines could be targeted. Therefore, in this article, we consider the potential values and uses of algorithms for characterising cDNAs amongst the collated helminth genomic information that encode secreted proteins, and methods for their selective isolation and cloning. We also review the variety of prokaryotic and eukaryotic cell expression systems that have been employed for the production and downstream purification of recombinant proteins in functionally active form, and provide an overview of the parameters that must be considered if these recombinant proteins are to be commercialised as vaccine therapeutics in humans and/or animals.

Animals↗

Mining genomic databases to identify novel hydrogen producers.

The realization that fossil fuel reserves are limited and their adverse effect on the environment has forced us to look into alternative sources of energy. Hydrogen is a strong contender as a future fuel. Biological hydrogen production ranges from 0.37 to 3.3 moles H(2) per mole of glucose and, considering the high theoretical values of production (4.0 moles H(2) per mole of glucose), it is worth exploring approaches to increase hydrogen yields. Screening the untapped microbial population is a promising possibility. Sequence analysis and pathway alignment of hydrogen metabolism in complete and incomplete genomes has led to the identification of potential hydrogen producers.

Bacteria↗

Mining genome databases to identify and understand new gene regulatory systems.

The availability of a large number of sequenced microbial genomes allows us to conduct systematic studies on microbial gene regulatory systems. Computational methods, using comparative genomics approaches, are powerful tools to understand their mechanisms and evolutionary history. Recent advances in computational methodology for uncovering transcriptional regulatory components and their interactions are discussed.

Computational Biology↗

Mining genomes: correlating tandem mass spectra of modified and unmodified peptides to sequences in nucleotide databases.

The correlation of uninterpreted tandem mass spectra of modified and unmodified peptides, produced under low-energy (10-50 eV) collision conditions, with nucleotide sequences is demonstrated. In this method nucleotide databases are translated in six reading frames, and the resulting amino acid sequences are searched "on the fly" to identify and fit linear sequences to the fragmentation patterns observed in the tandem mass spectra of peptides. A cross-correlation function is then used to provide a measurement of similarity between the mass-to-charge ratios for the fragment ions predicted by amino acid sequences translated from the nucleotide database and the fragment ions observed in the tandem mass spectrum. In general, a difference greater than 0.1 between the normalized cross-correlation functions for the first- and second-ranked search results indicates a successful match between sequence and spectrum. Measurements of the deviation from maximum similarity employing the spectral reconstruction method are made. The search method employing nucleotide databases is also demonstrated on the spectra of phosphorylated peptides. Specific sites of modification are identified even though no specific information relevant to sites of modification is contained in the character-based sequence information of nucleotide databases.

Amino Acid Sequence↗

Immuno-informatics: Mining genomes for vaccine components.

The complete genome sequences of more than 60 microbes have been completed in the past decade. Concurrently, a series of new informatics tools, designed to harness this new wealth of information, have been developed. Some of these new tools allow researchers to select regions of microbial genomes that trigger immune responses. These regions, termed epitopes, are ideal components of vaccines. When the new tools are used to search for epitopes, this search is usually coupled with in vitro screening methods; an approach that has been termed computational immunology or immuno-informatics. Researchers are now implementing these combined methods to scan genomic sequences for vaccine components. They are thereby expanding the number of different proteins that can be screened for vaccine development, while narrowing this search to those regions of the proteins that are extremely likely to induce an immune response. As the tools improve, it may soon be feasible to skip over many of the in vitro screening steps, moving directly from genome sequence to vaccine design. The present article reviews the work of several groups engaged in the development of immuno-informatics tools and illustrates the application of these tools to the process of vaccine discovery.

Algorithms↗

POCUS: mining genomic sequence annotation to predict disease genes.

Here we present POCUS (prioritization of candidate genes using statistics), a novel computational approach to prioritize candidate disease genes that is based on over-representation of functional annotation between loci for the same disease. We show that POCUS can provide high (up to 81-fold) enrichment of real disease genes in the candidate-gene shortlists it produces compared with the original large sets of positional candidates. In contrast to existing methods, POCUS can also suggest counterintuitive candidates.

Autistic Disorder↗

Bioinformatics pipeline for the systematic mining genomic and proteomic variation linked to rare diseases: The example of monogenic diabetes.

Monogenic diabetes is characterized as a group of diseases caused by rare variants in single genes. Like for other rare diseases, multiple genes have been linked to monogenic diabetes with different measures of pathogenicity, but the information on the genes and variants is not unified among different resources, making it challenging to process them informatically. We have developed an automated pipeline for collecting and harmonizing data on genetic variants linked to monogenic diabetes. Furthermore, we have translated variant genetic sequences into protein sequences accounting for all protein isoforms and their variants. This allows researchers to consolidate information on variant genes and proteins linked to monogenic diabetes and facilitates their study using proteomics or structural biology. Our open and flexible implementation using Jupyter notebooks enables tailoring and modifying the pipeline and its application to other rare diseases.

Humans↗