Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Vaccination against Neisseria meningitidis using three variants of the lipoprotein GNA1870.

Sepsis and meningitis caused by serogroup B meningococcus are devastating diseases of infants and young adults, which cannot yet be prevented by vaccination. By genome mining, we discovered GNA1870, a new surface-exposed lipoprotein of Neisseria meningitidis that induces high levels of bactericidal antibodies. The antigen is expressed by all strains of N. meningitidis tested. Sequencing of the gene in 71 strains representative of the genetic and geographic diversity of the N. meningitidis population, showed that the protein can be divided into three variants. Conservation within each variant ranges between 91.6 to 100%, while between the variants the conservation can be as low as 62.8%. The level of expression varies between strains, which can be classified as high, intermediate, and low expressors. Antibodies against a recombinant form of the protein elicit complement-mediated killing of the strains that carry the same variant and induce passive protection in the infant rat model. Bactericidal titers are highest against those strains expressing high yields of the protein; however, even the very low expressors are efficiently killed. The novel antigen is a top candidate for the development of a new vaccine against meningococcus.

Adult↗

Protective antibody responses elicited by a meningococcal outer membrane vesicle vaccine with overexpressed genome-derived neisserial antigen 1870.

Background. Meningococcal outer membrane vesicle (OMV) vaccines are efficacious in humans but have serosubtype-specific serum bactericidal antibody responses directed at the porin protein PorA and the potential for immune selection of PorA-escape mutants.Methods. We prepared an OMV vaccine from a Neisseria meningitidis strain engineered to overexpress genome-derived neisserial antigen (GNA) 1870, a lipoprotein discovered by genome mining that is being investigated for use in a vaccine.Results. Mice immunized with the modified GNA1870-OMV vaccine developed broader serum bactericidal antibody responses than control mice immunized with a recombinant GNA1870 protein vaccine or an OMV vaccine prepared from wild-type N. meningitidis or a combination of vaccines prepared from wild-type N. meningitidis and recombinant protein. Antiserum from mice immunized with the modified GNA1870-OMV vaccine also elicited greater deposition of human C3 complement on the surface of live N. meningitidis bacteria and greater passive protective activity against meningococcal bacteremia in infant rats. A N. meningitidis mutant with decreased expression of PorA was more susceptible to bactericidal activity of anti-GNA1870 antibodies.Conclusions. The modified GNA1870-OMV vaccine elicits broader protection against meningococcal disease than recombinant GNA1870 protein or conventional OMV vaccines and also has less risk of selection of PorA-escape mutants than a conventional OMV vaccine.

Animals↗

The Celera Discovery System.

The Celera Discovery System (CDS) is a web-accessible research workbench for mining genomic and related biological information. Users have access to the human and mouse genome sequences with annotation presented in summary form in BioMolecule Reports for genes, transcripts and proteins. Over 40 additional databases are available, including sequence, mapping, mutation, genetic variation, mRNA expression, protein structure, motif and classification data. Data are accessible by browsing reports, through a variety of interactive graphical viewers, and by advanced query capability provided by the LION SRS search engine. A growing number of sequence analysis tools are available, including sequence similarity, pattern searching, multiple sequence alignment and Hidden Markov Model search. A user workspace keeps track of queries and analyses. CDS is widely used by the academic research community and requires a subscription for access. The system and academic pricing information are available at http://cds.celera.com.

Animals↗

Biocontrol effect of a solid-state fermentation-derived extract mixture of Trichoderma asperellum on sunflower Sclerotinia rot and associated host defense responses.

Sclerotinia disease is a destructive fungal disease of sunflowers, soybeans, and other economically important crops, causing substantial yield loss and quality deterioration. Long-term reliance on dose-dependent broad-spectrum fungicides is constrained by resistance risks and potential environmental burdens, creating tension with the sustainability goal of "reducing pesticide use while improving efficacy." Here, we explore a Trichoderma spp.-based microbial disease management strategy. Whole-genome sequencing of Trichoderma asperellum TCS007 isolated from Antarctic marine sediments, coupled with genome mining, predicted diverse biosynthetic gene clusters putatively associated with siderophores, polyketides, nonribosomal peptides, and terpenoids; the corresponding metabolites are not chemically confirmed and require further validation. Using a solid-state fermentation workflow, we prepared a fermentation-derived extract mixture (TCS007-SSF-Ex). In vitro assays showed dose-dependent inhibition of Sclerotinia sclerotiorum by TCS007-SSF-Ex (EC50 = 1.252 mg/L), and microscopy revealed cellular damage-consistent changes, including organelle disruption and plasmolysis. Pathogen transcriptomic and metabolism-related analyses indicated broad perturbations in organelle biogenesis and metabolic processes, with significant alterations in pathways associated with succinate, D-glucose, and phenylacetate; these results are consistent with growth inhibition and reduced pathogenicity, but specific molecular targets and causal links remain to be validated. In vivo, under certain application conditions, triple applications increased APX activity (+492.5%) and β-1,3-glucanase activity (+419.6%). Collectively, this work supports a "pathogen suppression-host defense induction" framework and facilitates subsequent identification of active components and mechanistic validation.IMPORTANCESclerotinia diseases cause recurrent and economically important losses in oilseed crops, while long-term fungicide use is constrained by resistance risks and environmental burdens. Trichoderma-based biocontrol is a promising complementary strategy, yet evidence supporting metabolite-containing Trichoderma-derived preparations as immune elicitors remains less consolidated than that for living inoculants, and scalable production routes are still needed. Here, we examine an Antarctic marine sediment-derived strain, Trichoderma asperellum TCS007, and a solid-state fermentation (SSF)-derived extract mixture (TCS007-SSF-Ex) produced via solid-state fermentation. We combine in vitro antifungal assays, pathogen ultrastructural observations, and correlative omics analyses with in vivo measurements of sunflower defense enzymes (APX and β-1,3-glucanase) to evaluate a "pathogen suppression-host defense induction" framework. Our findings support the potential of SSF-derived Trichoderma metabolite mixtures for greener management of Sclerotinia disease and provide a foundation for future chemical identification of active components and mechanistic validation.

Ascomycota↗

Fighting infection using immunomodulatory agents.

The last decade has seen the emergence of immunomodulators as promising therapeutic agents in infectious diseases. A diverse array of recombinant, synthetic and natural immunomodulatory preparations for prophylaxis and treatment of various infections are available today. Some of these substances, such as granulocyte colony-stimulating factor (G-CSF), interferons, imiquimod and bacterial-derived preparations are already licensed for use in patients. Others including IL-12, various chemokines, synthetic cytosine phosphate-guanosine (CpG) oligodeoxynucleotides and glucans are being investigated extensively in clinical and preclinical studies. Immunomodulatory regimens offer an attractive approach as an adjunct modality for control of microbial diseases in the era of antibiotic resistance. Practical application of the advances in molecular biology, bioinformatics, genomic mining and high-throughput peptide synthesis should foster future discovery and development of novel immunomodulators contingent upon scientific evidence rather than dictates of discursive empiricism.

Adjuvants, Immunologic↗

Artificial Intelligence for Natural Products Discovery and Development.

Natural products (NPs) remain a cornerstone of modern drug discovery, offering stereochemical complexity and diverse bioactivities that precisely modulate therapeutic targets, refined through billions of years of evolution. However, their research has long been hindered by inefficient, empirical workflows, high resource consumption, structural complexity, and the "multicomponent, multi-target" nature of their mechanisms. The exponential growth of genomic, metabolomic, and spectral data has overwhelmed conventional analytical methods, exposing critical bottlenecks in handling high-dimensional, heterogeneous datasets that exceed human interpretive capacity. Artificial intelligence (AI) is emerging as a transformative paradigm to address these challenges, integrating multi-omics and chemical data to shift NP research from fragmented empiricism toward mechanism-driven, precision-oriented development. By leveraging deep learning architectures- including graph neural networks, Transformers, and diffusion-based generative models-AI enables systematic decoding of NP biosynthesis, automated structure elucidation, rational target identification, knowledge extraction from vast unstructured scientific literature, and de novo molecular design. This review comprehensively surveys recent advances in AI applications across the full NP discovery and development pipeline, encompassing genome mining, structure-based and ligand-based virtual screening, multimodal structural characterization, lead optimization, and biosynthetic pathway engineering. We further examine the emerging roles of protein-centric, molecule- centric, and multimodal foundation models, as well as large language models, in bridging genotype-to-chemotype gaps and unlocking unstructured scientific knowledge. Finally, we discuss critical challenges including data scarcity, representational limitations for complex stereochemistry, physical plausibility in generative models, and the urgent need for experimental validation, while outlining future directions toward autonomous experimentation, closed-loop optimization, and human-AI collaborative discovery.

Artificial intelligence↗

Recombinant semaphorin 6A-1 ectodomain inhibits in vivo growth factor and tumor cell line-induced angiogenesis.

The Semaphorins are a large family of transmembrane, GPI-anchored and secreted proteins that play an important role in neuronal and endothelial cell guidance. A human gene related to the class 6 Semaphorin family, Semaphorin 6A-1 (Sema 6A-1) was identified by homology-based genomic mining. Recent implication of Sema 3 family members in tumor angiogenesis and our expression analysis of Sema 6A-1 suggested that class 6 Semaphorin might effect tumor neovascularization. The mRNA expression of Sema 6A-1 was elevated in several renal tumor tissue samples relative to adjacent nontumor tissue samples from the same patient. Sema 6A-1 transcript was also expressed in the majority of renal clear cell carcinoma (RCC) cell lines and to a lesser extent in endothelial cells. To test the role of Sema 6A-1 in tumor angiogenesis, we engineered, expressed and purified the Sema 6A-1 soluble extracellular domain (Sema-ECD). The purified Sema-ECD was screened in a variety of endothelial cell-based assays both in vitro and in vivo. In vitro, Sema-ECD blocked VEGF-mediated endothelial cell migration. These effects were explained in part by our observation in endothelial cells that Sema-ECD inhibited VEGF-mediated Src, FAK and ERK phosphorylation. In vivo, mouse Matrigel assays demonstrated that the intraperitoneal administration of recombinant Sema-ECD inhibited both bFGF/VEGF and tumor cell line-induced neovascularization. These findings reveal a novel therapeutic utility for Sema 6A-1 (Sema-ECD) as an inhibitor of growth factor as well as tumor-induced angiogenesis.

Adenocarcinoma, Clear Cell↗

Mining proteases in the genome databases.

Protease data mining can take advantage both of the many specialist, Web-available databases that cover the genetic, protein and nucleic acid sequence information that is specific to a variety of organisms, and of a flexible, but defined, classification system. However, precomputed data, such as gene predictions, should be used with care. Unless there is definitive supporting information, ideally sequencing of a cDNA to show that the predictions are accurate, followed by expression and biochemical characterization of the predicted protein, the predicted gene and its product remains a possibility, rather than a certainty.

Animals↗

Database of repetitive elements in complete genomes and data mining using transcription factor binding sites.

Approximately 43% of the human genome is occupied by repetitive elements. Even more, around 51% of the rice genome is occupied by repetitive elements. The analysis presented here indicates that repetitive elements in complete genomes may have been very important in the evolutionary genomics. In this study, a database, called the Repeat Sequence Database, is first designed and implemented to store complete and comprehensive repetitive sequences. See http://rsdb.csie.ncu.edu.tw for more information. The database contains direct, inverted and palindromic repetitive sequences, and each repetitive sequence has a variable length ranging from seven to many hundred nucleotides. The repetitive sequences in the database are explored using a mathematical algorithm to mine rules on how combinations of individual binding sites are distributed among repetitive sequences in the database. Combinations of transcription factor binding sites in the repetitive sequences are obtained and then data mining techniques are applied to mine association rules from these combinations. The discovered associations are further pruned to remove insignificant associations and obtain a set of associations. The mined association rules facilitate efforts to identify gene classes regulated by similar mechanisms and accurately predict regulatory elements. Experiments are performed on several genomes including C. elegans, human chromosome 22, and yeast.

Algorithms↗

Crystallization data mining in structural genomics: using positive and negative results to optimize protein crystallization screens.

Recent efforts to collect and mine crystallization data from structural genomics (SG) consortia have led to the identification of minimal screens and novel screening strategies that can be used to streamline the crystallization process. Two groups, the Joint Center for Structural Genomics and the University of Toronto, carried out large-scale crystallization trials on different sets of bacterial targets (539, JCSG and 755, Toronto), using different sample processing and crystallization methods, and then analyzed their results to identify the smallest subset of conditions that would have crystallized the maximum number of protein targets. The JCSG Core Screen contains 67 conditions (from 480) while the Toronto Minimal Screen contains 6 (from 48). While the exact conditions included in the two screens do not overlap, the major precipitants of the conditions are similar and thus both screens can be used to determine if a protein has a natural propensity to crystallize. In addition, studies from other groups including the University of Queensland, the Mycobacterium tuberculosis SG group, the Southeast Collaboratory for SG, and the York Structural Biology Laboratory indicate that alternative crystallization strategies may be more successful at identifying initial crystallization conditions than typical sparse matrix screens. These minimal screens and alternative screening strategies are already being used to optimize the crystallization processes within large SG efforts. The differences between these results, however, demonstrate that additional studies which examine the influence of protein biophysical properties and sample preparation methods on crystal formation must also be carried out before more robust screens can be identified.

Chemistry Techniques, Analytical↗

Sight: automating genomic data-mining without programming skills.

SUMMARY: We created and tested Sight, a Java-based package that provides a user-friendly interface to generate and connect agents for automatic genomic data-mining for individual requirements without requiring programming skills from the user. AVAILABILITY: http://physiologie.uni-ulm.de//Seiten/Arbeitsgruppe/Jurkat-Rott/Jurkat-Rott.htm. The system does not require additional components and runs on IBM PCs under Windows (NT 4.0, 2000 and XP) or Linux (Phat 4.0 and Mandrake 9.0).

Algorithms↗

Transcriptome mining and comparative genomics reveal 36 putative novel marafivirus species and conserved evolution of the marafibox regulatory element.

BACKGROUND: Marafiviruses are plant-infecting RNA viruses associated with several economically important crops, but their genomic diversity remains incompletely characterized. OBJECTIVE: This study aimed to identify previously unrecognized marafivirus genomes and investigate their genomic features and evolutionary relationships. METHODS: Publicly available plant transcriptome datasets were systematically mined to detect marafivirus-like sequences. Recovered genomes were analyzed using comparative sequence analysis, phylogenetic reconstruction, and genome organization characterization. RESULTS: A total of 62 marafivirus-like genomes were recovered from 33 independent sources representing diverse plant hosts. Polyprotein-based comparative and phylogenetic analyses grouped these genomes into 36 lineages likely representing novel species. All newly identified viruses clustered within the Marafivirus clade. Genome organization analysis revealed conserved polyprotein architecture and widespread presence of the marafibox promoter element. Conservation of additional open reading frames among closely related isolates aided identification of potentially functional genes. CONCLUSION: These findings substantially expand the known diversity of marafiviruses and demonstrate the effectiveness of transcriptome mining for discovering previously unrecognized plant viruses.

Phylogeny↗

WormBase: methods for data mining and comparative genomics.

WormBase is a comprehensive repository for information on Caenorhabditis elegans and related nematodes. Although the primary web-based interface of WormBase (http:// www.wormbase.org/) is familiar to most C. elegans researchers, WormBase also offers powerful data-mining features for addressing questions of comparative genomics, genome structure, and evolution. In this chapter, we focus on data mining at WormBase through the use of flexible web interfaces, custom queries, and scripts. The intended audience includes users wishing to query the database beyond the confines of the web interface or fetch data en masse. No knowledge of programming is necessary or assumed, although users with intermediate skills in the Perl scripting language will be able to utilize additional data-mining approaches.

Animals↗

GeneMerge--post-genomic analysis, data mining, and hypothesis testing.

SUMMARY: GeneMerge is a web-based and standalone program written in PERL that returns a range of functional and genomic data for a given set of study genes and provides statistical rank scores for over-representation of particular functions or categories in the data set. Functional or categorical data of all kinds can be analyzed with GeneMerge, facilitating regulatory and metabolic pathway analysis, tests of population genetic hypotheses, cross-experiment comparisons, and tests of chromosomal clustering, among others. GeneMerge can perform analyses on a wide variety of genomic data quickly and easily and facilitates both data mining and hypothesis testing. AVAILABILITY: GeneMerge is available free of charge for academic use over the web and for download from: http://www.oeb.harvard.edu/hartl/lab/publications/GeneMerge.html.

Algorithms↗

The NEIBank project for ocular genomics: data-mining gene expression in human and rodent eye tissues.

NEIBank is a project to gather and organize genomic resources for eye research. The first phase of this project covers the construction and sequence analysis of cDNA libraries from human and animal model eye tissues to develop an overview of the repertoire of genes expressed in the eye and a resource of cDNA clones for further studies. The sequence data are grouped and identified using the tools of bioinformatics and the results are displayed through a web site where they can be interrogated by keyword search, chromosome location, by Blast (sequence comparison) or by alignment on completed genomes. Many novel proteins and novel splice forms of known genes have already emerged from analysis of the accumulating data. This review provides an overview of the current state of the database for human eye tissues, with specific comparisons to some parallel data from mouse and rat, and with illustrative examples of the kinds of insights and discoveries these data can produce. One of the major themes that emerges is that at the molecular level human eye tissues have significant differences from those of rodents, encompassing species specific genes, alternative splice forms and great variation in levels of gene expression. These point to specific adaptations and mechanisms in the human eye and emphasize that care needs to be taken in the application of appropriate animal model systems.

Amino Acid Sequence↗

Mining the structural genomics pipeline: identification of protein properties that affect high-throughput experimental analysis.

Structural genomics projects represent major undertakings that will change our understanding of proteins. They generate unique datasets that, for the first time, present a standardized view of proteins in terms of their physical and chemical properties. By analyzing these datasets here, we are able to discover correlations between a protein's characteristics and its progress through each stage of the structural genomics pipeline, from cloning, expression, purification, and ultimately to structural determination. First, we use tree-based analyses (decision trees and random forest algorithms) to discover the most significant protein features that influence a protein's amenability to high-throughput experimentation. Based on this, we identify potential bottlenecks in various stages of the structural genomics process through specialized "pipeline schematics". We find that the properties of a protein that are most significant are: (i.) whether it is conserved across many organisms; (ii). the percentage composition of charged residues; (iii). the occurrence of hydrophobic patches; (iv). the number of binding partners it has; and (v). its length. Conversely, a number of other properties that might have been thought to be important, such as nuclear localization signals, are not significant. Thus, using our tree-based analyses, we are able to identify combinations of features that best differentiate the small group of proteins for which a structure has been determined from all the currently selected targets. This information may prove useful in optimizing high-throughput experimentation. Further information is available from http://mining.nesg.org/.

Algorithms↗

Mining the human genome for new health therapies.

The completion of the Human Genome Project heralds advances in determining the foundations of disease and in developing new therapeutic treatments. Tests already exist for the detection of some genetic abnormalities that can cause disease, and more are being developed. In the future, pharmacogenetics will be used to tailor treatment to specific patients.

Female↗

Http://C. elegans: mining the functional genomic landscape.

Caenorhabditis elegans is a powerful animal model for the study of functional genomics. The completed and well-annotated DNA sequence is available and a systematic study of gene function by RNA-interference-mediated knockdown of every gene is in progress. Full-genome DNA microarrays and DNA chips can be used to determine expression changes at different stages of development and in different mutant backgrounds, and a protein-interaction map based on the yeast two-hybrid approach is in progress. These high-capacity approaches to studying gene function will provide new insights into invertebrate and vertebrate biology.

Animals↗