Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Mining the human genome using microarrays of open reading frames.

To test the hypothesis that the human genome project will uncover many genes not previously discovered by sequencing of expressed sequence tags (ESTs), we designed and produced a set of microarrays using probes based on open reading frames (ORFs) in 350 Mb of finished and draft human sequence. Our approach aims to identify all genes directly from genomic sequence by querying gene expression. We analysed genomic sequence with a suite of ORF prediction programs, selected approximately one ORF per gene, amplified the ORFs from genomic DNA and arrayed the amplicons onto treated glass slides. Of the first 10,000 arrayed ORFs, 31% are completely novel and 29% are similar, but not identical, to sequences in public databases. Approximately one-half of these are expressed in the tissues we queried by microarray. Subsequent verification by other techniques confirmed expression of several of the novel genes. Expressed sequence tags (ESTs) have yielded vast amounts of data, but our results indicate that many genes in the human genome will only be found by genomic sequencing.

Cell Line↗

Mining the Plasmodium genome database to define organellar function: what does the apicoplast do?

Apicomplexan species constitute a diverse group of parasitic protozoa, which are responsible for a wide range of diseases in many organisms. Despite differences in the diseases they cause, these parasites share an underlying biology, from the genetic controls used to differentiate through the complex parasite life cycle, to the basic biochemical pathways employed for intracellular survival, to the distinctive cell biology necessary for host cell attachment and invasion. Different parasites lend themselves to the study of different aspects of parasite biology: Eimeria for biochemical studies, Toxoplasma for molecular genetic and cell biological investigation, etc. The Plasmodium falciparum Genome Project contributes the first large-scale genomic sequence for an apicomplexan parasite. The Plasmodium Genome Database (http://PlasmoDB.org) has been designed to permit individual investigators to ask their own questions, even prior to formal release of the reference P. falciparum genome sequence. As a case in point, PlasmoDB has been exploited to identify metabolic pathways associated with the apicomplexan plastid, or 'apicoplast' - an essential organelle derived by secondary endosymbiosis of an alga, and retention of the algal plastid.

Animals↗

Data mining the Arabidopsis genome reveals fifteen 14-3-3 genes. Expression is demonstrated for two out of five novel genes.

In plants, 14-3-3 proteins are key regulators of primary metabolism and membrane transport. Although the current dogma states that 14-3-3 isoforms are not very specific with regard to target proteins, recent data suggest that the specificity may be high. Therefore, identification and characterization of all 14-3-3 (GF14) isoforms in the model plant Arabidopsis are important. Using the information now available from The Arabidopsis Information Resource, we found three new GF14 genes. The potential expression of these three genes, and of two additional novel GF14 genes (Rosenquist et al., 2000), in leaves, roots, and flowers was examined using reverse transcriptase-polymerase chain reaction and cDNA library polymerase chain reaction screening. Under normal growth conditions, two of these genes were found to be transcribed. These genes were named grf11and grf12, and the corresponding new 14-3-3 isoforms were named GF14omicron and GF14iota, respectively. The gene coding for GF14omicron was expressed in leaves, roots, and flowers, whereas the gene coding for GF14iota was only expressed in flowers. Gene structures and relationships between all members of the GF14 gene family were deduced from data available through The Arabidopsis Information Resource. The data clearly support the theory that two 14-3-3 genes were present when eudicotyledons diverged from monocotyledons. In total, there are 15 14-3-3 genes (grfs 1-15) in Arabidopsis, of which 12 (grfs 1-12) now have been shown to be expressed.

14-3-3 Proteins↗

Antigen mining with iterative genome screens identifies novel diagnostics for the Mycobacterium tuberculosis complex.

The definition of antigens for the diagnosis of human and bovine tuberculosis is a research priority. If diagnosis is to be used alongside Mycobacterium bovis BCG-based vaccination regimens, it will be necessary to have reagents that allow the discrimination of infected and vaccinated animals. A list of 42 potential M. bovis-specific antigens was prepared by comparative analysis of the genomes of M. bovis, M. avium subsp. avium, M. avium subsp. paratuberculosis, and Streptomyces coelicolor. Potential antigens were tested by applying them in a high-throughput peptide-based screening system to M. bovis-infected and BCG-vaccinated cattle and to cattle without prior exposure to M. bovis. A response hierarchy of antigens was established by comparing responses in infected animals. Three antigens (Mb2555, Mb2890, and Mb3895) were selected for further study, as they were strongly recognized in experimentally infected animals but with low or no frequency in BCG-vaccinated and naïve cows. Interestingly, all three antigens were recognized in animals vaccinated against Johne's disease, suggesting the presences of epitopes cross-reacting with M. avium subsp. paratuberculosis antigens. Eight peptides from the three antigens studied in detail were identified as immunodominant and were characterized in terms of major histocompatibility complex class II restriction element usage and shown to be restricted through both DR and DQ molecules. Reasons for antigenic cross-reactivity with M. avium subsp. paratuberculosis and refinement of the in silico strategy to predict such cross-reactivity from the primary protein sequence will be discussed. Evaluation of the peptides identified from the three dominant antigens by use of larger field studies is now a priority.

Amino Acid Sequence↗

CGH-Profiler: data mining based on genomic aberration profiles.

BACKGROUND: CGH-Profiler is a program that supports the analysis of genomic aberrations measured by Comparative Genomic Hybridisation (CGH). Comparative genomic hybridisation (CGH) is a well-established, molecular cytogenetic method that allows the detection of chromosomal imbalances in entire genomes. This technique is widely used in routine molecular diagnostics. Typically, chromosomal imbalances are described in a complex syntax based on the International Standard for Cytogenetic Nomenclature (ISCN). This semantic description of chromosomal imbalances hinders a large-scale statistical analysis across different experiments, e.g. for finding aberration patterns associated with a particular disease type or state. RESULTS: CGH-Profiler circumvents the semantic ISCN description by importing data from different CGH system vendors and by directly transferring the data into a table format that is readily accessible for subsequent statistical analysis. CGH-profiler comes with different consistency checks, calculates various statistics and automatically assigns a median copy number ratio to each chromosomal band. Import of CGH profiles from different CGH system vendors is already supported; its extension to other systems can be readily achieved through Perl scripts.CGH profiler can also be used to analyse comparative expressed sequence hybridisation (CESH) data. CESH reveals gene expression patterns according to chromosomal locations in a similar manner as CGH detects chromosomal imbalances. CONCLUSION: CGH-Profiler is a useful tool for processing of CGH and CESH data.

Chromosome Aberrations↗

Mining the probiotic genome: advanced strategies, enhanced benefits, perceived obstacles.

Recent advances in DNA sequencing has made it possible to accurately decipher the entire genetic complement of a probiotic bacterium. Increases in sequencing capabilities have been enhanced through improved computer software that can annotate, or identify, the majority of genes encoded by the sequence. The availability of annotated genome sequence will be important in defining the capabilities of the individual strains of probiotic bacteria. It will also form the platform for microarray and proteomic technologies that allow real-time analysis of RNA and protein expression in the bacterial cell. Investigation of probiotic organisms with these new and potentially powerful tools will facilitate the development of the bacteria as therapeutic agents, and provide the mechanisms to produce advanced probiotic strains. This paper addresses the core technologies in the rapidly growing area of genomics, and their application to the molecular characterisation of probiotic bacteria and host-microbe interactions.

Databases, Genetic↗

Mining the mammalian genome for artiodactyl systematics.

A total of 7,806 nucleotide positions derived from one mitochondrial and eight nuclear DNA segments were used to provide a robust phylogeny for members of the order Artiodactyla. Twenty-four artiodactyl and two cetacean species were included, and the horse (order Perissodactyla) was used as the outgroup. Limited rate heterogeneity was observed among the nuclear genes. The partition homogeneity tests indicated no conflicting signal among the nuclear genes fragments, so the sequence data were analyzed together and as separate loci. Analyses based on the individual nuclear DNA fragments and on 34 unique indels all produced phylogenies largely congruent with the topology from the combined data set. In sharp contrast to the nuclear DNA data, the mtDNA cytochrome b sequence data showed high levels of homoplasy, failed to produce a robust phylogeny, and were remarkably sensitive to taxon sampling. The nuclear DNA data clearly support the paraphyletic nature of the Artiodactyla. Additionally, the family Suidae is diphyletic, and the nonruminating pigs and peccaries (Suiformes) were the most basal cetartiodactyl group. The morphologically derived Ruminantia was always monophyletic; within this group, all taxa with paired bony structures on their skulls clustered together. The nuclear DNA data suggest that the Antilocaprinae account for a unique evolutionary lineage, the Cervidae and Bovidae are sister taxa, and the Giraffidae are more primitive.

Animals↗

Bioinformatics: use in bacterial vaccine discovery.

Bioinformatics has now become a common laboratory name for groups studying genomic sequences. It is composed of many different, yet interrelated scientific fields such as genomics, proteomics, and transcriptional profiling. The availability of complete genomic sequences, especially prokaryotic organisms, allows one to rapidly identify, analyze, and clone genes of interest. For bacterial vaccine discovery, one can "mine" the genomic sequence for potential surface targets using various algorithms, characterize these gene targets, and produce primers for cloning, all before one enters the wet laboratory. This review will focus on various genomic mining tools/algorithms available for predicting open reading frames and their associated annotation (if known), physical and functional characterization, and cellular localization. Finally, examples are given of how all of this is being used for the identification of potential bacterial vaccine candidates.

Animals↗

General and robust sample preparation strategies for cryo-EM studies of CRISPR-Cas9 and Cas12 enzymes.

Cas9 and Cas12 are RNA-guided DNA endonucleases derived from prokaryotic CRISPR-Cas adaptive immune systems that have been repurposed as versatile genome-engineering tools. Computational mining of genomes and metagenomes has expanded the diversity of Cas9 and Cas12 enzymes that can be used to develop versatile, orthogonal molecular toolboxes. Structural information is pivotal to uncovering the precise molecular mechanisms of newly discovered Cas enzymes and providing a foundation for their application in genome editing. In this chapter, we describe detailed protocols for the preparation of Cas9 and Cas12 enzymes for cryo-electron microscopy. These methods will enable fast and robust structural determination of newly discovered Cas9 and Cas12 enzymes, which will enhance the understanding of diverse CRISPR-Cas effectors and provide a molecular framework for expanding CRISPR-based genome-editing technologies.

Cryoelectron Microscopy↗

Genomic comparison using data mining techniques based on a possibilistic fuzzy sets model.

Current copiousness of genomic information stored in biological databases [Mar Albà, M., Lee, M., Pearl, D., Shepherd, F.M.G., Martin, A.J., Orengo, N., Kellam, C.A., 2001. P. VIDA: a virus database system for the organisation of virus genome open reading frames. Nuleic Acids Res. 133-136] makes ultimately feasible the proposal for an application of knowledge management aimed to discover general rules in subcellular phenomena. The goal of this work is primarily to discover relationships between genes by microarray analysis. The tools exploited come from clustering techniques and are mainly based on Knowledge Discovery in Databases (KDD) concepts [Fayyad, U., Piatetsky-Shapiro, G., Smyth, P., 1996. From data mining to knowledge discovery in databases. AI Magazine 17(3), 37-54]. Starting from a data set, each element can be represented by a characteristic matrix, which sums up all data attributes. In this case data mining is oriented to perform a Pattern Recognition of related sequences, hidden in databases [Hand, D.J., Nicholas, A., 2005. Heard finding groups in gene expression data. J. Biomed. Biotechnol. 215-225]. Following a bottom up approach, the next refinement is to compare retrieved data to gather similar features, by dedicated clustering algorithms [Kaufman, L., Rousseeuw, P.J., 1990. Finding groups in data. An Introduction to Cluster Analysis. John Wiley & Sons, New York; Forman, G., Zhang, B., 2000. Distributed Data clustering can be efficient and exact HP. Laboratories Palo Alto HPL-2000, p. 158], driven by fuzzy logic, allowing us to perceive by intuition a common denominator for various genomic families and to anticipate likely future developments.

Algorithms↗

BBP: Brucella genome annotation with literature mining and curation.

BACKGROUND: Brucella species are Gram-negative, facultative intracellular bacteria that cause brucellosis in humans and animals. Sequences of four Brucella genomes have been published, and various Brucella gene and genome data and analysis resources exist. A web gateway to integrate these resources will greatly facilitate Brucella research. Brucella genome data in current databases is largely derived from computational analysis without experimental validation typically found in peer-reviewed publications. It is partially due to the lack of a literature mining and curation system able to efficiently incorporate the large amount of literature data into genome annotation. It is further hypothesized that literature-based Brucella gene annotation would increase understanding of complicated Brucella pathogenesis mechanisms. RESULTS: The Brucella Bioinformatics Portal (BBP) is developed to integrate existing Brucella genome data and analysis tools with literature mining and curation. The BBP InterBru database and Brucella Genome Browser allow users to search and analyze genes of 4 currently available Brucella genomes and link to more than 20 existing databases and analysis programs. Brucella literature publications in PubMed are extracted and can be searched by a TextPresso-powered natural language processing method, a MeSH browser, a keywords search, and an automatic literature update service. To efficiently annotate Brucella genes using the large amount of literature publications, a literature mining and curation system coined Limix is developed to integrate computational literature mining methods with a PubSearch-powered manual curation and management system. The Limix system is used to quickly find and confirm 107 Brucella gene mutations including 75 genes shown to be essential for Brucella virulence. The 75 genes are further clustered using COG. In addition, 62 Brucella genetic interactions are extracted from literature publications. These results make possible more comprehensive investigation of Brucella pathogenesis. Other BBP features include publication email alert service, Brucella researchers' contact database, and discussion forum. CONCLUSION: BBP is a gateway for Brucella researchers to search, analyze, and curate Brucella genome data originated from public databases and literature. Brucella gene mutations and genetic interactions are annotated using Limix leading to better understanding of Brucella pathogenesis.

Algorithms↗

CAGEcleaner: reducing genomic redundancy in gene cluster mining.

SUMMARY: Mining homologous biosynthetic gene clusters (BGCs) typically involves searching colocalised genes against large genomic databases. However, the high degree of genomic redundancy in these databases often propagates into the resulting hit sets, complicating downstream analyses and visualization. To address this challenge, we present CAGEcleaner, a Python-based pipeline with auxiliary bash scripts designed to reduce redundancy in gene cluster hit sets by dereplicating the genomes that host these hits. CAGEcleaner integrates seamlessly with widely used gene cluster mining tools, such as cblaster and CAGECAT, enabling efficient filtering and streamlining BGC discovery workflows. AVAILABILITY AND IMPLEMENTATION: Source code and documentation is hosted at GitHub (https://github.com/LucoDevro/CAGEcleaner) and Zenodo (https://doi.org/10.5281/zenodo.14726119) under an MIT license. For accessibility, CAGEcleaner is installable from Bioconda (https://anaconda.org/bioconda/cagecleaner) and PyPi (https://pypi.org/project/cagecleaner/), and is also available as a Docker image from DockerHub (https://hub.docker.com/r/lucodevro/cagecleaner).

Software↗

Monitoring of bacteria in acid mine environments by reverse sample genome probing.

A variety of microorganisms can exist in acid mine drainage (AMD) environments, although their contribution to AMD problems is unclear. Environmental strains of Thiobacillus ferrooxidans and Thiobacillus acidophilus were purified by repeated plating and single-colony isolation on iron salts and tetrathionate media, respectively. Thiobacillus thiooxidans was enriched on sulfur-containing media. For the isolation of Leptospirillum ferrooxidans, iron salts and pyrite media were inoculated with environmental samples. However, L. ferrooxidans was never recovered on solid media. Denatured chromosomal DNAs from type and (or) isolated strains of T. ferrooxidans, T. acidophilus, T. thiooxidans, and L. ferrooxidans were spotted on a master filter for their detection in a variety of samples by reverse sample genome probing (RSGP). Analysis of enrichments of environmental samples by RSGP indicated that ferrous sulfate medium enriched T. ferrooxidans strains, whereas all thiobacilli grew in sulfur medium, T. thiooxidans strains being dominant. Enrichment in glucose medium followed by transfer to tetrathionate medium resulted in the selection of T. acidophilus strains. DNA was also extracted directly (without enrichment) from cells recovered from AMD water or sediments, and was analyzed by RSGP to describe the communities present. Strains showing homology with T. ferrooxidans and T. acidophilus were found to be major community components. Strains showing homology with T. thiooxidans were a minor community component, whereas strains showing homology with L. ferrooxidans were not detected.

Acids↗

The next generation of literature analysis: integration of genomic analysis into text mining.

Text-mining systems are indispensable tools to reduce the increasing flux of information in scientific literature to topics pertinent to a particular interest in focus. Most of the scientific literature is published as unstructured free text, complicating the development of data processing tools, which rely on structured information. To overcome the problems of free text analysis, structured, hand-curated information derived from literature is integrated in text-mining systems to improve precision and recall. In this paper several text-mining approaches are reviewed and the next step in development of text-mining systems, which is based on a concept of multiple lines of evidence, is described: results from literature analysis are combined with evidence from experiments and genome analysis to improve the accuracy of results and to generate additional knowledge beyond what is known solely from literature.

Abstracting and Indexing↗