Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Microbial pathogen genomes - new strategies for identifying therapeutic and vaccine targets.

Efficient mining of genomic sequence information from multiple pathogens for therapeutic and vaccine targets requires efficient tools. Fortunately, robust methods applicable to whole genomes have been developed and applied in the past few years to identify genes essential for growth or virulence and to detect potential vaccine targets. Successful approaches to identify potential therapeutic targets include a variety of ingenious uses of nearly random transposon insertions, more directed methods such as antisense and insertion-duplication mutagenesis, and expression profiling facilitated by microarrays. Vaccine targets have been identified by gene fusion and expression experiments to discover gene products that are immunogenic in humans or animal models. All genome-wide methods require focused secondary assays to validate the findings, but these genomic methods excel at reducing to a manageable number the genes to be examined further. This editorial reviews the latest developments in genome-wide target identification tools.

Anti-Infective Agents↗

Onogenomics: dissecting cancer through genome research.

The Oncogenomics meeting focused on bioinformatics, molecular pathways and global gene expression profiles relating to cancer. Several sessions were devoted to updating the audience on the latest status of the human genome project. Future directions will focus on mining the genome for new information about the genetic code in humans. Proteomics is becoming a useful tool for helping to understand the structure and function of proteins and their partners, which will, in turn, enable us to more rationally use proteins as targets for therapy.

Journal Article↗

Identification of Gasz, an evolutionarily conserved gene expressed exclusively in germ cells and encoding a protein with four ankyrin repeats, a sterile-alpha motif, and a basic leucine zipper.

To discover causes of infertility and potential contraceptive targets, we used in silico subtraction and genomic database mining to identify conserved genes with germ cell-specific expression. In silico subtraction identified an expressed sequence tag (EST) present exclusively in a newborn mouse ovary library. The full-length cDNA sequence corresponding to this EST encodes a novel protein containing four ankyrin (ANK) repeats, a sterile-alpha motif (SAM), and a putative basic leucine zipper (bZIP) domain. Northern blot and semiquantitative RT-PCR analyses demonstrated that the mRNA is exclusively expressed in the mouse testis and ovary. The expression sites were localized by in situ hybridization to pachytene spermatocytes in the testis and oocytes in the ovary. Immunohistochemistry showed that the novel protein is localized to the cytoplasm in pachytene spermatocytes and early spermatids, oocytes at all stages of oogenesis, and in early preimplantation embryos. Based on its germ cell-specific expression and the presence of ANK, SAM, and basic leucine zipper domains, we have termed this novel protein GASZ. The mouse Gasz gene, which consists of 13 exons and spans 60 kb, is located on chromosome 6 between the Wnt2 and cystic fibrosis transmembrane conductance regulator (Cftr) genes. Using genomic database mining, orthologous genes encoding GASZ were identified in the rat, cow, baboon, chimpanzee, and human. Phylogenetic analyses reveal that the GASZ proteins are highly conserved among these species. Human and mouse GASZ proteins share 85.3% amino acid identity, and human and chimpanzee GASZ proteins differ by only 3 out of 475 amino acids. In humans, the GASZ gene resides on chromosome 7 and is similarly composed of 13 exons. Because both ANK repeats and the SAM domain function as protein-protein interaction modules that mediate signal transduction cascades in some systems, GASZ may represent an important cytoplasmic signal transducer that mediates protein-protein interactions during germ cell maturation in both males and females and during preimplantation embryogenesis.

Adaptor Proteins, Signal Transducing↗

Mining the draft human genome.

Now that the draft human genome sequence is available, everyone wants to be able to use it. However, we have perhaps become complacent about our ability to turn new genomes into lists of genes. The higher volume of data associated with a larger genome is accompanied by a much greater increase in complexity. We need to appreciate both the scale of the challenge of vertebrate genome analysis and the limitations of current gene prediction methods and understanding.

Animals↗

Mining the Giardia lamblia genome for new cyst wall proteins.

The Giardia lamblia cyst wall (CW), which is required for survival outside the host and infection, is a primitive extracellular matrix. Because of the importance of the CW, we queried the Giardia Genome Project Database with the coding sequences of the only two known CW proteins, which are cysteine-rich and contain leucine-rich repeats (LRRs). We identified five new LRR-containing proteins, of which only one (CWP3) is up-regulated during encystation and incorporated into the cyst wall. Sequence comparison with CWP1 and -2 revealed conservation within the LRRs and the 44-amino-acid N-flanking region, although CWP3 is more divergent. Interestingly, all 14 cysteine residues of CWP3 are positionally conserved with CWP1 and -2. During encystation, C-terminal epitope-tagged CWP3 was transported to the wall of water-resistant cysts via the novel regulated secretory pathway in encystation-secretory vesicles (ESVs). Deletion analysis revealed that the four LRRs are each essential to target CWP3 to the ESVs and cyst wall. In a deletion of the most C-terminal region, fewer ESVs were stained in encysting cells, and there was no staining in cysts. In contrast, deletion of the 44 amino acids between the signal sequence and the LRRs or the region just C-terminal to the LRRs only decreased the number of cells with CWP3 targeting to ESVs and cyst wall by approximately 50%. Our studies indicate that virtually every portion of the CWP3 protein is needed for efficient targeting to the regulated secretory pathway and incorporation into the cyst wall. Further, these data demonstrate the power of genomics in combination with rigorous functional analyses to verify annotation.

Amino Acid Sequence↗

yMGV: a database for visualization and data mining of published genome-wide yeast expression data.

The yeast Microarray Global Viewer (yMGV) is an on-line database providing a synthetic view of the transcriptional expression profiles of Saccharomyces cerevisiae genes in most of the published expression datasets. yMGV displays a one-screen graphical representation of gene expression variations for each published genome-wide experiment, allowing quick retrieval of experimental conditions affecting expression of this gene. yMGV also provides tools to isolate groups of genes sharing similar transcription profiles in a defined subset of experiments. Additionally, yMGV furnishes a set of statistical tools for critical assessment of published data. We therefore believe that yMGV is an efficient tool that affords a quick and comprehensive overview of microarray data and generates new gene classifications. As of 20 March 2001 the yMGV database contains 6 000 000 measurements, representing genome-wide expression comparisons of 932 experiments from 39 microarray publications. The yMGV interface is available at http://transcriptome.ens.fr/ymgv/.

Computational Biology↗

SNP-RFLPing: restriction enzyme mining for SNPs in genomes.

BACKGROUND: The restriction fragment length polymorphism (RFLP) is a common laboratory method for the genotyping of single nucleotide polymorphisms (SNPs). Here, we describe a web-based software, named SNP-RFLPing, which provides the restriction enzyme for RFLP assays on a batch of SNPs and genes from the human, rat, and mouse genomes. RESULTS: Three user-friendly inputs are included: 1) NCBI dbSNP "rs" or "ss" IDs; 2) NCBI Entrez gene ID and HUGO gene name; 3) any formats of SNP-in-sequence, are allowed to perform the SNP-RFLPing assay. These inputs are auto-programmed to SNP-containing sequences and their complementary sequences for the selection of restriction enzymes. All SNPs with available RFLP restriction enzymes of each input genes are provided even if many SNPs exist. The SNP-RFLPing analysis provides the SNP contig position, heterozygosity, function, protein residue, and amino acid position for cSNPs, as well as commercial and non-commercial restriction enzymes. CONCLUSION: This web-based software solves the input format problems in similar softwares and greatly simplifies the procedure for providing the RFLP enzyme. Mixed free forms of input data are friendly to users who perform the SNP-RFLPing assay. SNP-RFLPing offers a time-saving application for association studies in personalized medicine and is freely available at http://bio.kuas.edu.tw/snp-rflp/.

Animals↗

Mining the Arabidopsis thaliana genome for highly-divergent seven transmembrane receptors.

To identify divergent seven-transmembrane receptor (7TMR) candidates from the Arabidopsis thaliana genome, multiple protein classification methods were combined, including both alignment-based and alignment-free classifiers. This resolved problems in optimally training individual classifiers using limited and divergent samples, and increased stringency for candidate proteins. We identified 394 proteins as 7TMR candidates and highlighted 54 with corresponding expression patterns for further investigation.

Arabidopsis↗

Genomic research and data-mining technology: implications for personal privacy and informed consent.

This essay examines issues involving personal privacy and informed consent that arise at the intersection of information and communication technology (ICT) and population genomics research. I begin by briefly examining the ethical, legal, and social implications (ELSI) program requirements that were established to guide researchers working on the Human Genome Project (HGP). Next I consider a case illustration involving deCODE Genetics, a privately owned genetic company in Iceland, which raises some ethical concerns that are not clearly addressed in the current ELSI guidelines. The deCODE case also illustrates some ways in which an ICT technique known as data mining has both aided and posed special challenges for researchers working in the field of population genomics. On the one hand, data-mining tools have greatly assisted researchers in mapping the human genome and in identifying certain "disease genes" common in specific populations (which, in turn, has accelerated the process of finding cures for diseases tha affect those populations). On the other hand, this technology has significantly threatened the privacy of research subjects participating in population genomics studies, who may, unwittingly, contribute to the construction of new groups (based on arbitrary and non-obvious patterns and statistical correlations) that put those subjects at risk for discrimination and stigmatization. In the final section of this paper I examine some ways in which the use of data mining in the context of population genomics research poses a critical challenge for the principle of informed consent, which traditionally has played a central role in protecting the privacy interests of research subjects participating in epidemiological studies.

Computational Biology↗

The complete human olfactory subgenome.

Olfactory receptors likely constitute the largest gene superfamily in the vertebrate genome. Here we present the nearly complete human olfactory subgenome elucidated by mining the genome draft with gene discovery algorithms. Over 900 olfactory receptor genes and pseudogenes (ORs) were identified, two-thirds of which were not annotated previously. The number of extrapolated ORs is in good agreement with previous theoretical predictions. The sequence of at least 63% of the ORs is disrupted by what appears to be a random process of pseudogene formation. ORs constitute 17 gene families, 4 of which contain more than 100 members each. "Fish-like" Class I ORs, previously considered a relic in higher tetrapods, constitute as much as 10% of the human repertoire, all in one large cluster on chromosome 11. Their lower pseudogene fraction suggests a functional significance. ORs are disposed on all human chromosomes except 20 and Y, and nearly 80% are found in clusters of 6-138 genes. A novel comparative cluster analysis was used to trace the evolutionary path that may have led to OR proliferation and diversification throughout the genome. The results of this analysis suggest the following genome expansion history: first, the generation of a "tetrapod-specific" Class II OR cluster on chromosome 11 by local duplication, then a single-step duplication of this cluster to chromosome 1, and finally an avalanche of duplication events out of chromosome 1 to most other chromosomes. The results of the data mining and characterization of ORs can be accessed at the Human Olfactory Receptor Data Exploratorium Web site (http://bioinfo.weizmann.ac.il/HORDE).

Chromosome Mapping↗

Sequence- and structure-based protein function prediction from genomic information.

Existing functional annotation transfer is fraught with inaccuracies that may hinder forward interpretation and mining of genomic data. Hand-curation of the annotation placed into databases is not practical. In lieu of experimental evidence, computational biological approaches offer high-throughput tools to predict function accurately; however, these methods are still notably deficient in defining and describing the complexity of protein function. Enriching genomic sequences obtained from sequencing efforts and expression array methods with protein function information and classification will be an efficient first step for incorporating genomic data into drug discovery programs.

Computational Biology↗

A microsatellite map of white clover.

The white clover ( Trifolium repens) nuclear genome (n = 2x = 16) is an important yet under-characterised genetic environment. We have developed simple sequence repeat (SSR) genetic markers for the white clover genome by mining an expressed sequence tag (EST) database and by isolation from enriched genomic libraries. A total of 2,086 EST-derived SSRs (EST-SSRs) were identified among 26,480 database accessions. Evaluation of 792 EST-SSR primer pairs resulted in 566 usable EST-SSRs. Of these, 335 polymorphic EST-SSRs, used in concert with 30 genomic SSRs, detected 493 loci in the white clover genome using 92 F1 progeny from a pair cross between two highly heterozygous genotypes--364/7 and 6525/5. Map length, as estimated using the joinmap algorithm, was 1,144 cM and spanned all 16 homologues. The R (red leaf) locus was mapped to linkage group B1 and is tightly linked to the microsatellite locus prs318c. The eight homoeologous pairs of linkage groups within the white clover genome were identified using 96 homoeologous loci. Segregation distortion was detected in four areas (groups A1, D1, D2 and H2). Marker locus density varied among and within linkage groups. This is the first time EST-SSRs have been used to build a whole-genome functional map and to describe subgenome organisation in an allopolyploid species, and T. repens is the only Trifolieae species to date to be mapped exclusively with SSRs. This gene-based microsatellite map will enable the resolution of quantitative traits into Mendelian characters, the characterisation of syntenic relationships with other genomes and acceleration of white clover improvement programmes.

Chromosome Mapping↗

PLATCOM: a Platform for Computational Comparative Genomics.

MOTIVATION: As more whole genome sequences become available, comparing multiple genomes at the sequence level can provide insight into new biological discovery. However, there are significant challenges for genome comparison. The challenge includes requirement for computational resources owing to the large volume of genome data. More importantly, since the choice of genomes to be compared is entirely subjective, there are too many choices for genome comparison. For these reasons, there is pressing need for bioinformatics systems for comparing multiple genomes where users can choose genomes to be compared freely. RESULTS: PLATCOM (Platform for Computational Comparative Genomics) is an integrated system for the comparative analysis of multiple genomes. The system is built on several public databases and a suite of genome analysis applications are provided as exemplary genome data mining tools over these internal databases. Researchers are able to visually investigate genomic sequence similarities, conserved gene neighborhoods, conserved metabolic pathways and putative gene fusion events among a set of selected multiple genomes. AVAILABILITY: http://platcom.informatics.indiana.edu/platcom

Chromosome Mapping↗

GeneCards: a novel functional genomics compendium with automated data mining and query reformulation support.

MOTIVATION: Modern biology is shifting from the 'one gene one postdoc' approach to genomic analyses that include the simultaneous monitoring of thousands of genes. The importance of efficient access to concise and integrated biomedical information to support data analysis and decision making is therefore increasing rapidly, in both academic and industrial research. However, knowledge discovery in the widely scattered resources relevant for biomedical research is often a cumbersome and non-trivial task, one that requires a significant amount of training and effort. RESULTS: To develop a model for a new type of topic-specific overview resource that provides efficient access to distributed information, we designed a database called 'GeneCards'. It is a freely accessible Web resource that offers one hypertext 'card' for each of the more than 7000 human genes that currently have an approved gene symbol published by the HUGO/GDB nomenclature committee. The presented information aims at giving immediate insight into current knowledge about the respective gene, including a focus on its functions in health and disease. It is compiled by Perl scripts that automatically extract relevant information from several databases, including SWISS-PROT, OMIM, Genatlas and GDB. Analyses of the interactions of users with the Web interface of GeneCards triggered development of easy-to-scan displays optimized for human browsing. Also, we developed algorithms that offer 'ready-to-click' query reformulation support, to facilitate information retrieval and exploration. Many of the long-term users turn to GeneCards to quickly access information about the function of very large sets of genes, for example in the realm of large-scale expression studies using 'DNA chip' technology or two-dimensional protein electrophoresis. AVAILABILITY: Freely available at http://bioinformatics.weizmann.ac.il/cards/ CONTACT: cards@bioinformatics.weizmann.ac.il

Algorithms↗

Unraveling the diversity, function, and virus-host interactions of archaeal proviruses.

Archaea, the third domain of life, play critical roles in global biogeochemical cycles. However, archaeal proviruses integrated into host genomes remain largely unexplored. To bridge this gap, we conducted a large-scale mining of genomes spanning all presently known 21 archaeal phyla for their proviruses. We identified 770 archaeal proviruses across 12 archaeal phyla and 84 families, which clustered into 655 viral operational taxonomic units (vOTUs). Among these, 86.1% of the vOTUs were novel at the species level, and 69.3% could not be classified at the family level, substantially expanding the known diversity of archaeal viruses. Additionally, phylogenomic analysis supported the proposal of 16 putative novel viral families, further extending the current taxonomy landscape of archaeal viruses. Notably, 21.8% of the identified proviruses were predicted to adopt a lytic lifestyle, suggesting that these proviruses may retain the capacity to enter the lytic cycle under appropriate conditions. Host prediction indicated only 14 out of the 655 vOTUs might have potential across-lineage infection abilities. We detected 63 anti-defense genes encoded by 61 provirus genomes, such as anti-CRISPR and anti-RM, suggesting an ongoing evolutionary arms race between hosts and proviruses. However, only 10 auxiliary metabolic genes (AMGs) were identified, suggesting a limited impact of proviruses in the modulation of host metabolism through AMGs. This study establishes a systematic global genomic atlas of archaeal proviruses, advancing our understanding of their distribution and diversity while providing a foundation for future research into how proviruses regulate archaeal metabolism and ecosystem functioning.

anti-defense system↗

Homologs of the alpha- and beta-subunits of mammalian brain platelet-activating factor acetylhydrolase Ib in the Drosophila melanogaster genome.

The mammalian intracellular brain platelet-activating factor acetylhydrolase, implicated in the development of cerebral cortex, is a member of the phospholipase A2 superfamily. It is made up of a homodimer of the 45 kDa LIS1 protein (a product of the causative gene for type I lissencephaly) and a pair of homologous 26-kDa alpha-subunits which account for all the catalytic activity. LIS1 is hypothesized to regulate nuclear movement in migrating neurons through interactions with the cytoskeleton, while the alpha-subunits, whose structure is known, contain a trypsin-like triad within the framework of a unique tertiary fold. The physiological significance of the association of the two types of subunits is not known. In an effort to better understand the function of the complex we turned to genomic data mining in search of related proteins in lower eukaryotes. We found that the Drosophila melanogaster genome contains homologs of both alpha- and beta-subunits, and we cloned both genes. The alpha-subunit homolog has been overexpressed, purified and crystallized. It lacks two of the three active-site residues and, consequently, is catalytically inactive against PAF-AH (Ib) substrates. Our study shows that the beta-subunit homolog is highly conserved from Drosophila to mammals and is able to interact with the mammalian alpha-subunits but is unable to interact with the Drosophila alpha-subunit. Proteins 2000;39:1-8.

1-Alkyl-2-acetylglycerophosphocholine Esterase↗

G-language Genome Analysis Environment: a workbench for nucleotide sequence data mining.

SUMMARY: G-language Genome Analysis Environment (G-language GAE) is an open source generic software package aimed for higher efficiency in bioinformatics analysis. G-language GAE has an interface as a set of Perl libraries for software development, and a graphical user interface for easy manipulation. Both Windows and Linux versions are available. AVAILABILITY: From http://www.g-language.org/ under GNU General Public License. CD-ROMs are distributed freely in major conferences.

Database Management Systems↗