Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

The use of functional genomics in C. elegans for studying human development and disease.

The 100 Mb Caenorhabditis elegans genome sequence is the first animal genome to be sequenced in its entirety. Many reverse-genetics tools have been developed to mine the genome sequence and to facilitate the jump between the identification of a gene sequence and the understanding of its function. Here we discuss how C. elegans can contribute to understanding of the function of genes involved in human development and disease.

Animals↗

Genome-wide analysis of core cell cycle genes in the unicellular green alga Ostreococcus tauri.

The cell cycle has been extensively studied in various organisms, and the recent access to an overwhelming amount of genomic data has given birth to a new integrated approach called comparative genomics. Comparing the cell cycle across species shows that its regulation is evolutionarily conserved; the best-known example is the pivotal role of cyclin-dependent kinases in all the eukaryotic lineages hitherto investigated. Interestingly, the molecular network associated with the activity of the CDK-cyclin complexes is also evolutionarily conserved, thus, defining a core cell cycle set of genes together with lineage-specific adaptations. In this paper, we describe the core cell cycle genes of Ostreococcus tauri, the smallest free-living eukaryotic cell having a minimal cellular organization with a nucleus, a single chloroplast, and only one mitochondrion. This unicellular marine green alga, which has diverged at the base of the green lineage, shows the minimal yet complete set of core cell cycle genes described to date. It has only one homolog of CDKA, CDKB, CDKD, cyclin A, cyclin B, cyclin D, cyclin H, Cks, Rb, E2F, DP, DEL, Cdc25, and Wee1. We have also added the APC and SCF E3 ligases to the core cell cycle gene set. We discuss the potential of genome-wide analysis in the identification of divergent orthologs of cell cycle genes in different lineages by mining the genomes of evolutionarily important and strategic organisms.

Algal Proteins↗

TBP-associated factors in Arabidopsis.

Initiation of transcription mediated by RNA polymerase II requires a number of transcription factors among which TFIID is the major core promoter recognition factor. TFIID is composed of highly conserved factors which include the TATA-binding protein (TBP) and about 14 TBP-associated factors (TAFs). Since TAFs play important roles in transcription they have been extensively studied in organisms like yeast, Drosophila and human. Surprisingly, TAFs have been poorly characterized in plants. With the completion of the Arabidopsis genome sequence, it is possible to search for TAFs, since many of them have conserved amino acid sequences. Mining the genome of Arabidopsis for TAFs resulted in the identification of 18 putative Arabidopsis TAFs (AtTAFs). We have analyzed their protein structure and their genomic localisation. Expression profiling by RT-PCR showed that these TAFs are expressed in all parts of the plant which is in agreement with their general role in transcription. These analyses in combination with their evolutionary conservation with TAFs of other organisms are discussed.

Arabidopsis↗

Genomic analysis of secretion systems.

Secretion of proteins into the extracellular environment is important to almost all bacteria, and in particular mediates interactions between pathogenic or symbiotic bacteria with their eukaryotic hosts. The accumulation of bacterial genome sequence data in the past few years has provided great insights into the distribution and function of these secretion systems. Three systems are responsible for secretion of proteins across the bacterial cytoplasmic membrane: Sec, SRP and Tat. Many novel examples of systems for transport across the Gram-negative bacterial cell envelope have been discovered through genome sequencing and surveys, including many novel type III secretion systems and autotransporters. Similarly, genomic data mining has revealed many new potential secretion substrates and identified unsuspected domains in secretion-associated proteins. Interestingly, genomic analyses have also hinted at the existence of a dedicated protein secretion system in Gram-positive bacteria, targeting members of the WXG100/ESAT-6 family of proteins, and have revealed an unexpectedly wide distribution of sortase-driven protein-targeting systems.

Bacteria↗

Alu insertion polymorphisms for the study of human genomic diversity.

Genomic database mining has been a very useful aid in the identification and retrieval of recently integrated Alu elements from the human genome. We analyzed Alu elements retrieved from the GenBank database and identified two new Alu subfamilies, Alu Yb9 and Alu Yc2, and further characterized Yc1 subfamily members. Some members of each of the three subfamilies have inserted in the human genome so recently that about a one-third of the analyzed elements are polymorphic for the presence/absence of the Alu repeat in diverse human populations. These newly identified Alu insertion polymorphisms will serve as identical-by-descent genetic markers for the study of human evolution and forensics. Three previously classified Alu Y elements linked with disease belong to the Yc1 subfamily, supporting the retroposition potential of this subfamily and demonstrating that the Alu Y subfamily currently has a very low amplification rate in the human genome.

Alu Elements↗

Phylogenetic and structural relationships of the PR5 gene family reveal an ancient multigene family conserved in plants and select animal taxa.

Pathogenesis-related group 5 (PR5) plant proteins include thaumatin, osmotin, and related proteins, many of which have antimicrobial activity. The recent discovery of PR5-like (PR5-L) sequences in nematodes and insects raises questions about their evolutionary relationships. Using complete plant genome data and discovery of multiple insect PR5-L sequences, phylogenetic comparisons among plants and animals were performed. All PR5/PR5-L protein sequences were mined from genome data of a member of each of two main angiosperm groups-the eudicots (Arabidoposis thaliana) and the monocots (Oryza sativa)-and from the Caenorhabditis nematode (C. elegans and C. briggsase). Insect PR5-L sequences were mined from EST databases and GenBank submissions from four insect orders: Coleoptera (Diaprepes abbreviatus and Biphyllus lunatus), Orthoptera (Schistocerca gregaria), Hymenoptera (Lysiphlebus testaceipes), and Hemiptera (Toxoptera citricida). Parsimony and Bayesian phylogenetic analyses showed that the PR5 family is paraphyletic in plants, likely arising from 10 genes in a common ancestor to monocots and eudicots. After evolutionary divergence of monocots and eudicots, PR5 genes increased asymmetrically among the 10 clades. Insects and nematodes contain multiple sequences (seven PR5-Ls in nematodes and at least three in some insects) all related to the same plant clade, with nematode and insect sequences separating as two clades. Protein structural homology modeling showed strong similarity among animal and plant PR5/PR5-Ls, with divergence only in surface-exposed loops. Sequence and structural conservation among PR5/PR5-Ls suggests an important and conserved role throughout the evolutionary divergence of the diverse organisms from which they reside.

Amino Acid Sequence↗

Knowledge of the Bacillus subtilis genome: impacts on fundamental science and biotechnology.

The advent of genomics has greatly influenced fundamental and applied microbiology. This has become paradigmatic in the case of Bacillus subtilis, a primary model bacterium for research and biotechnology. Indeed, mining its genome has provided more fruitful information than classical approaches would have yielded in a longer period of time. Through advanced analysis of its genome and transcriptome, fundamental discoveries dealing with the informational architecture of the B. subtilis chromosome, as well as with the elucidation of its pathway-level regulation of gene expression, have been achieved. The possibility of performing a complete metabolic manipulation of the secretory pathway of Bacillus is promising important biotechnological fallouts. Similar emphasis exists for the possibility of controlling the cell in the formation of biofilms with specific physical and chemical characteristics. At the theoretical level, the new concept of genetic superinformation has been formulated and its analytical approach implemented, while the understanding of the minimal genetic requirements for the existence of a reproducing bacterial cell is being tackled. In summary, the impact of the B. subtilis genome has philosophically revolutionised the way that basic knowledge is translated into applied microbiology and biotechnology, making this bacterium the workhorse of post-genomic microbiology.

Bacillus subtilis↗

GeneFarm, structural and functional annotation of Arabidopsis gene and protein families by a network of experts.

Genomic projects heavily depend on genome annotations and are limited by the current deficiencies in the published predictions of gene structure and function. It follows that, improved annotation will allow better data mining of genomes, and more secure planning and design of experiments. The purpose of the GeneFarm project is to obtain homogeneous, reliable, documented and traceable annotations for Arabidopsis nuclear genes and gene products, and to enter them into an added-value database. This re-annotation project is being performed exhaustively on every member of each gene family. Performing a family-wide annotation makes the task easier and more efficient than a gene-by-gene approach since many features obtained for one gene can be extrapolated to some or all the other genes of a family. A complete annotation procedure based on the most efficient prediction tools available is being used by 16 partner laboratories, each contributing annotated families from its field of expertise. A database, named GeneFarm, and an associated user-friendly interface to query the annotations have been developed. More than 3000 genes distributed over 300 families have been annotated and are available at http://genoplante-info.infobiogen.fr/Genefarm/. Furthermore, collaboration with the Swiss Institute of Bioinformatics is underway to integrate the GeneFarm data into the protein knowledgebase Swiss-Prot.

Arabidopsis↗

GeneLynx mouse: integrated portal to the mouse genome.

GeneLynx Mouse is a meta-database providing an extensive collection of hyperlinks to mouse gene-specific information in diverse databases available via the Internet. The GeneLynx project is based on the simple notion that given any gene-specific identifier (e.g., accession number, gene name, text, or sequence), scientists should be able to access a single location that provides a set of links to all the publicly available information pertinent to the specified gene. The recent climax in the mouse genome and RIKEN cDNA sequencing projects provided the data necessary for the development of a gene-centric mouse information portal based on the GeneLynx ideals. Clusters of RIKEN cDNA sequences were used to define the initial set of mouse genes. Like its human counterpart, GeneLynx Mouse is designed as an extensible relational database with an intuitive and user-friendly Web interface. Data is automatically extracted from diverse resources, using appropriate approaches to maximize the coverage. To promote cross-database interoperability, an indexing utility is provided to facilitate the establishment of hyperlinks in external databases. As a result of the integration of the human and mouse systems, GeneLynx now serves as a powerful comparative genomics data mining resource. GeneLynx Mouse can be freely accessed at http://mouse.genelynx.org.

Animals↗

Presence and expression of hydrogenase specific C-terminal endopeptidases in cyanobacteria.

BACKGROUND: Hydrogenases catalyze the simplest of all chemical reactions: the reduction of protons to molecular hydrogen or vice versa. Cyanobacteria can express an uptake, a bidirectional or both NiFe-hydrogenases. Maturation of those depends on accessory proteins encoded by hyp-genes. The last maturation step involves the cleavage of a ca. 30 amino acid long peptide from the large subunit by a C-terminal endopeptidase. Until know, nothing is known about the maturation of cyanobacterial NiFe-hydrogenases. The availability of three complete cyanobacterial genome sequences from strains with either only the uptake (Nostoc punctiforme ATCC 29133/PCC 73102), only the bidirectional (Synechocystis PCC 6803) or both NiFe-hydrogenases (Anabaena PCC 7120) prompted us to mine these genomes for hydrogenase maturation related genes. In this communication we focus on the presence and the expression of the NiFe-hydrogenases and the corresponding C-terminal endopeptidases, in the three strains mentioned above. RESULTS: We identified genes encoding putative cyanobacterial hydrogenase specific C-terminal endopeptidases in all analyzed cyanobacterial genomes. The genes are not part of any known hydrogenase related gene cluster. The derived amino acid sequences show only low similarity (28-41%) to the well-analyzed hydrogenase specific C-terminal endopeptidase HybD from Escherichia coli, the crystal structure of which is known. However, computational secondary and tertiary structure modeling revealed the presence of conserved structural patterns around the highly conserved active site. Gene expression analysis shows that the endopeptidase encoding genes are expressed under both nitrogen-fixing and non-nitrogen-fixing conditions. CONCLUSION: Anabaena PCC 7120 possesses two NiFe-hydrogenases and two hydrogenase specific C-terminal endopeptidases but only one set of hyp-genes. Thus, in contrast to the Hyp-proteins, the C-terminal endopeptidases are the only known hydrogenase maturation factors that are specific. Therefore, in accordance with previous nomenclature, we propose the gene names hoxW and hupW for the bidirectional and uptake hydrogenase processing endopeptidases, respectively. Due to their constitutive expression we expect that, at least in cyanobacteria, the endopeptidases take over multiple functions.

Amino Acid Sequence↗

Regulatory networks: linking microarray data to systems biology.

Gene regulation and aging are intrinsically linked and these links often reach directly to transcription factors and their actions in gene regulation. However, it is very difficult to follow all the individual directions such factors can affect. Therefore, the opposite approach became more popular recently, i.e. observing the endpoints of all these actions. Microarrays are the preferred technology to monitor large-scale changes in transcripts across whole genomes. The trade-off for being able to survey whole genome transcriptomes is that the results are mere observations, which do not directly reveal the underlying mechanisms that represent the real link to transcription factors and their actions. Fortunately, a combination of knowledge mining (including but not restricted to literature mining) with genomics analyses can be harnessed to elucidate at least some of the regulatory networks orchestrating the transcriptional changes observed by microarray experiments. Thus, a considerable part of the functional system structure of cells and organisms can be revealed, which is a pivotal prerequisite for any meaningful systems biology approach towards aging related phenotypes.

Animals↗

Specialized microbial databases for inductive exploration of microbial genome sequences.

BACKGROUND: The enormous amount of genome sequence data asks for user-oriented databases to manage sequences and annotations. Queries must include search tools permitting function identification through exploration of related objects. METHODS: The GenoList package for collecting and mining microbial genome databases has been rewritten using MySQL as the database management system. Functions that were not available in MySQL, such as nested subquery, have been implemented. RESULTS: Inductive reasoning in the study of genomes starts from "islands of knowledge", centered around genes with some known background. With this concept of "neighborhood" in mind, a modified version of the GenoList structure has been used for organizing sequence data from prokaryotic genomes of particular interest in China. GenoChore http://bioinfo.hku.hk/genochore.html, a set of 17 specialized end-user-oriented microbial databases (including one instance of Microsporidia, Encephalitozoon cuniculi, a member of Eukarya) has been made publicly available. These databases allow the user to browse genome sequence and annotation data using standard queries. In addition they provide a weekly update of searches against the world-wide protein sequences data libraries, allowing one to monitor annotation updates on genes of interest. Finally, they allow users to search for patterns in DNA or protein sequences, taking into account a clustering of genes into formal operons, as well as providing extra facilities to query sequences using predefined sequence patterns. CONCLUSION: This growing set of specialized microbial databases organize data created by the first Chinese bacterial genome programs (ThermaList, Thermoanaerobacter tencongensis, LeptoList, with two different genomes of Leptospira interrogans and SepiList, Staphylococcus epidermidis) associated to related organisms for comparison.

Algorithms↗

Diversity and selection in sorghum: simultaneous analyses using simple sequence repeats.

Although molecular markers and DNA sequence data are now available for many crop species, our ability to identify genetic variation associated with functional or adaptive diversity is still limited. In this study, our aim was to quantify and characterize diversity in a panel of cultivated and wild sorghums (Sorghum bicolor), establish genetic relationships, and, simultaneously, identify selection signals that might be associated with sorghum domestication. We assayed 98 simple sequence repeat (SSR) loci distributed throughout the genome in a panel of 104 accessions comprising 73 landraces (i.e., cultivated lines) and 31 wild sorghums. Evaluation of SSR polymorphisms indicated that landraces retained 86% of the diversity observed in the wild sorghums. The landraces and wilds were moderately differentiated (F st=0.13), but there was little evidence of population differentiation among racial groups of cultivated sorghums (F st=0.06). Neighbor-joining analysis showed that wild sorghums generally formed a distinct group, and about half the landraces tended to cluster by race. Overall, bootstrap support was low, indicating a history of gene flow among the various cultivated types or recent common ancestry. Statistical methods (Ewens-Watterson test for allele excess, lnRH, and F st) for identifying genomic regions with patterns of variation consistent with selection gave significant results for 11 loci (approx. 15% of the SSRs used in the final analysis). Interestingly, seven of these loci mapped in or near genomic regions associated with domestication-related QTLs (i.e., shattering, seed weight, and rhizomatousness). We anticipate that such population genetics-based statistical approaches will be useful for re-evaluating extant SSR data for mining interesting genomic regions from germplasm collections.

Cluster Analysis↗

The gut's hidden arsenal: A genomics-guided atlas of class II bacteriocins.

Unmodified class II bacteriocins promise precision antimicrobials that spare bystander microbes. Zhang and colleagues introduce IIBacFinder, a genomics-guided pipeline that detects precursor and context genes with a curated pHMM library, infers leader-peptide cleavage, and triages candidates by meta-omics signals. The authors apply it across bacterial genomes, including an atlas of ∼280,000 human-gut genomes, and recover a vast reservoir of narrow-spectrum peptides and prioritize gut-resident candidates for synthesis. Of the 26 synthesized, 16 display activity in vitro, largely via membrane perturbation and with additive effects alongside vancomycin, while ex vivo assays show minimal compositional disruption of fecal communities compared with antibiotic controls. These results position unmodified class II bacteriocins as tractable, microbiome-sparing agents and illustrate how genome-scale mining coupled to meta-omics can bridge sequence to function in complex ecosystems.

Bacteriocins↗

Accelerating natural product discovery, characterization and engineering by biofoundries.

Covering: From early developments to the presentNatural product (NP) discovery is increasingly constrained by low-throughput screening, repeated rediscovery, and challenges in scaling genome mining-guided validation workflows. This highlight examines how automated biofoundries are accelerating NP discovery, characterization, and engineering through integrated design-build-test-learn (DBTL) cycles. We discuss recent advances in phenotype-first and genome-first discovery strategies enabled by robotics, high-throughput pathway reconstitution, and automated screening platforms. We further highlight emerging technologies, including cell-free biosynthesis, automated culturomics, programmable chassis engineering, and AI-assisted workflow orchestration, that may enable increasingly autonomous biofoundries for scalable exploration of NP chemical space and therapeutic discovery.

Journal Article↗

Bioinformatics, target discovery and the pharmaceutical/biotechnology industry.

With the first draft of the human genome now available a directed genome-wide mining strategy is being implemented by many pharmaceutical and biotechnology companies in order to identify novel members of the most therapeutically relevant target families. At the same time there is an increasing amount of annotation relevant to the human genome sequence entering into the public domain. The ability to identify protein families on a genome-wide scale can only be done at speed by using high-throughput computational approaches. This review describes many of the latest algorithmic developments in this field and shows how they can be best put to use for target identification and prioritization.

Algorithms↗

Alteration of host cell phenotype by Theileria annulata and Theileria parva: mining for manipulators in the parasite genomes.

The apicomplexan parasites Theileria annulata and Theileria parva cause severe lymphoproliferative disorders in cattle. Disease pathogenesis is linked to the ability of the parasite to transform the infected host cell (leukocyte) and induce uncontrolled proliferation. It is known that transformation involves parasite dependent perturbation of leukocyte signal transduction pathways that regulate apoptosis, division and gene expression, and there is evidence for the translocation of Theileria DNA binding proteins to the host cell nucleus. However, the parasite factors responsible for the inhibition of host cell apoptosis, or induction of host cell proliferation are unknown. The recent derivation of the complete genome sequence for both T. annulata and T. parva has provided a wealth of information that can be searched to identify molecules with the potential to subvert host cell regulatory pathways. This review summarizes current knowledge of the mechanisms used by Theileria parasites to transform the host cell, and highlights recent work that has mined the Theileria genomes to identify candidate manipulators of host cell phenotype.

Amino Acid Sequence↗

Cell envelope stress response in Bacillus licheniformis: integrating comparative genomics, transcriptional profiling, and regulon mining to decipher a complex regulatory network.

The envelope is an essential structure of the bacterial cell, and maintaining its integrity is a prerequisite for survival. To ensure proper function, transmembrane signal-transducing systems, such as two-component systems (TCS) and extracytoplasmic function (ECF) sigma factors, closely monitor its condition and respond to harmful perturbations. Both systems consist of a transmembrane sensor protein (histidine kinase or anti-sigma factor, respectively) and a corresponding cytoplasmic transcriptional regulator (response regulator or sigma factor, respectively) that mediates the cellular response through differential gene expression. The regulatory network of the cell envelope stress response is well studied in the gram-positive model organism Bacillus subtilis. It consists of at least two ECF sigma factors and four two-component systems. In this study, we describe the corresponding network in a close relative, Bacillus licheniformis. Based on sequence homology, domain architecture, and genomic context, we identified five TCS and eight ECF sigma factors as potential candidate regulatory systems mediating cell envelope stress response in this organism. We characterized the corresponding regulatory network by comparative transcriptomics and regulon mining as an initial screening tool. Subsequent in-depth transcriptional profiling was applied to define the inducer specificity of each identified cell envelope stress sensor. A total of three TCS and seven ECF sigma factors were shown to be induced by cell envelope stress in B. licheniformis. We noted a number of significant differences, indicative of a regulatory divergence between the two Bacillus species, in addition to the expected overlap in the respective responses.

Adaptation, Physiological↗