Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “bioinformatic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Hunting TPR Domains Using Kleisli.

We have two objectives. First, we want to build a system for detecting tetratricopeptide repeats in protein sequences. Second, we want to demonstrate how the general bioinformatics database integration system called Kleisli can help build such a system easily. We achieve these two objectives by showing that short and clear programs can be written in Kleisli, using its high-level query language CPL, to build a TPR domain hunter by integrating WU-BLAST2.0, HMMER, Entrez, and PFAM.

Journal Article↗

The ontology of the gene ontology.

The rapidly increasing wealth of genomic data has driven the development of tools to assist in the task of representing and processing information about genes, their products and their functions. One of the most important of these tools is the Gene Ontology (GO), which is being developed in tandem with work on a variety of bioinformatics databases. An examination of the structure of GO, however, reveals a number of problems, which we believe can be resolved by taking account of certain organizing principles drawn from philosophical ontology. We shall explore the results of applying such principles to GO with a view to improving GO's consistency and coherence and thus its future applicability in the automated processing of biological data.

Computational Biology↗

Assessment of virulence-factor activity relationships (VFARs) for waterborne diseases.

Virulence-factor activity relationship (VFAR) is a concept that was developed as a way to relate the architectural and biochemical components of a microorganism to its potential to cause human disease. Development of these relationships requires specialised bioinformatics databases that do not exist at present. A pilot-scale VFAR database was designed for three different waterborne organisms: Escherichia coli, Norovirus and Cryptosporidium, to evaluate VFAR relationships. For the web-based database, each organism has separate pages containing virulence genes, occurrence genes, primer sets and probes, taxonomy, outbreaks, and serotype/species/genogroup/genotype. As the database continues to grow, it will be possible to relate the occurrence and prevalence of certain genes in various microorganisms to outbreak data and, subsequently, to establish the utility of using a combination of specific genes as markers of virulence and in establishing virulence-factor activity relationships (VFARs). The database and the VFARs established will be of use to the regulatory community as a way to assist with prioritising those organisms, which need to be regulated.

Animals↗

[Molecular cloning and functional analysis of STGC3 - a novel gene on chromosome 3p21].

BACKGROUND & OBJECTIVE: A locus of loss of heterozygosity (LOH) with high frequency has been found on chromosome 3p21 in nasopharyngeal carcinoma (NPC). On the basis of our former research, this study was designed to clone and analyze a novel NPC-associated gene at this locus. METHODS: The full-length cDNA sequence of this gene was obtained by plasmid cDNA sequencing and RACE,and analyzed by bioinformatics. The pEGFP-C2/STGC3 fusion mammalian expression vector was constructed, and transfected into COS7 and CNE2 cell lines mediated by lipofectin to analyze the subcellular localization of gene expressing proteins. The expression of STGC3 was detected in normal tissues and tumor cell lines by Northern blot. RESULTS: An 1 271 bp full-length cDNA sequence of gene,which had no obvious homology with other known genes in bioinformatic databases,was obtained. This gene,named STGC3 (GenBank accession number:AY078383), localized on chromosome 3p21, and encodes a protein consisting of 146 amino acids. Protein localization analysis under fluorescent microscope indicated that STGC3 fusion protein was distributed in nucleus and cytoplasm 24-48 hours after transfection. MTE(TM)Array2 Northern blot analysis showed STGC3 expressed in both normal tissues and tumor cells,while its expression down-regulated in many tumor cell lines, such as Burkitt's lymphoma cell line Daudi. CONCLUSIONS: STGC3 is a novel gene, and down-regulated in NPC and many other tumor cell lines. STGC3 fusion protein distributes in cytoplasm and nucleus.

Amino Acid Sequence↗

In silico identification of breast cancer genes by combined multiple high throughput analyses.

Publicly available human genomic sequence data provide an unprecedented opportunity for researchers to decode the functionality of human genome. Such information is extremely valuable in cancer prevention diagnosis and treatment. Cancer Genome Anatomy Project (CGAP) and Gene Expression Omnibus (GEO) are two bioinformatic infrastructures for studying functional genomics. The goal of this study is to explore the feasibility of incorporating the Internet-available bioinformatic databases to discover human breast cancer-related genes. Several tools including the Gene Finder, Virtual Northern (vNorthern) and SAGE digital gene expression displayer (DGED) were used to analyze differential gene expression between benign and malignant breast tissues. A pilot study was performed using both EST and SAGE vNorthern to analyze the expression of a panel of known genes, including high abundance genes beta-actin and G3PDH, low abundance genes BRCA1 and p53, tissue-specific genes CEA and PSA and two breast cancer-related genes Her2/neu and MUC1. We found a high expression of beta-actin and G3PDH and a low expression of BRCA1 and p53 across different types of tissues as well as a tissue-specific expression of CEA in colon and PSA in prostate. A further analysis of 30 known breast cancer-related genes in breast cancer tissues by vNorthern demonstrated a high expression of oncogenes and low expression of tumor suppressor genes. An open-end analysis of two pools of breast cancer and benign breast tissue libraries by SAGE DGED produced 53 differentially expressed genes according to the screening criteria of a >five-fold difference and p<0.01. Further analysis by EST vNorthern and virtual microarray analysis reduced the candidate genes to six, with four down-regulated genes, ANXA1, CAV1, KRT5 and MMP7, and two up-regulated genes, ERBB2 and G1P3 in breast cancer. These findings were validated by a real-time RT-PCR analysis in eight paired human breast cancer tissue samples. We conclude that the combined multiple high throughput analyses is an effective data mining strategy in cancer gene identification. This approach may improve the usage of public available genomic data through strategic data mining of high throughput analysis.

Blotting, Northern↗

[In silicon cloning of HV126, a novel human gene related to multi drug resistance in leukemia].

OBJECTIVE: To find the novel gene related to the multi-drug resistance in leukemia and explore the molecular mechanism of multi-drug resistance. METHODS: The subtracted HL-60/VCR cDNA library was generated through the suppression subtractive hybridization using the wild HL-60 cells' cDNA as target and HL-60/ ATRA cells' as driver. A novel expression sequence tag (EST) sequence, which differentially expressed in HL-60/ ATRA cell, was screened by cDNA chip. Then a novel human gene, HV126 was assembled by the EST assembly tools. Bioinformatical databases and softwares were used to analyze and predict its function. Reverse transcription-PCR (RT-PCR) was used to detect the expression of HV-126 gene in leukemia cells before and after chemotherapy. RESULTS: The full open reading frames (ORFs) of the novel EST assembled by overlapping dbEST sequences included a 1991 bp nucleic sequence, which was named HV126. The deduced amino acid sequence consisted of 365 amino acids. The sequence of the novel gene exhibited 43% homology to a known gene, which is a possible member of the death domain-flood family implicated in apoptosis and inflammation. The expression of HV126 was proved to be related to the drug sensitivity in leukemia cells by RT-PCR. CONCLUSION: HV126, the novel gene, might have roles in regulating multi-drug resistance in leukemia. Further laboratory research should be done on cloning and making clear the gene function.

Antineoplastic Agents↗

Support vector machines for novel class detection in Bioinformatics.

Novelty detection techniques might be a promising way of dealing with high-dimensional classification problems in Bioinformatics. We present preliminary results of the use of a one-class support vector machine approach to detect novel classes in two Bioinformatics databases. The results are compatible with theory and inspire further investigation.

Artificial Intelligence↗

Integrating medical and genomic data: a successful example for rare diseases.

The recent advances on genomics and proteomics research bring up a significant grow on the information that is publicly available. However, navigating through genetic and bioinformatics databases can be a too complex and unproductive task for a primary care physician. In this paper we present diseasecard, a web portal for rare disease that provides transparently to the user a virtually integration of distributed and heterogeneous information.

Computational Biology↗

MetaBasis: a web-based database containing metadata on software tools and databases in the field of bioinformatics.

UNLABELLED: We have developed an integrated web-based relational database information system, which offers an extensive search functionality of validated entries containing available bioinformatics computing resources. This system, called MetaBasis, aims to provide the bioinformatics community, and especially newcomers to the field, with easy access to reliable bioinformatics databases and tools. MetaBasis is focused on non-commercial and open-source software tools. AVAILABILITY: http://metabasis.bioacademy.gr/

Computational Biology↗

The NCI/CIT microArray database (mAdb) system - bioinformatics for the management and analysis of Affymetrix and spotted gene expression microarrays.

A scalable, modular, enterprise-level system for both microarray databasing and analysis over the Internet has been developed over the past four years by the National Cancer Institute's Center for Cancer Research in collaboration with NIH's Center for Information Technology. This completely Web-based system, called mAdb (for microArray database), is currently supporting over 810 registered users and collaborators at NIH and contains over 22,000 microarray experiments, making it one of the largest collections of microarray data in existence. In addition, the mAdb system has been ported for the Netherlands Cancer Institute, the Genome Institute of Singapore, and the CDC. This system has been used for a wide variety of scientific experiments spanning the range from cancer to studies of early development, and for human, mouse, rat, yeast, and numerous microbial organisms.

Animals↗

[Bioinformatics and GenEnv database in biological risk management].

Identification and molecular typing of environmental isolates by molecular techniques requires knowledge of the genetic characteristics of the microbe species being examined. The introduction of automated sequences has greatly speeded up the entire sequencing process as well as improved the accuracy of the collected information. Bioinformatics tools have become indispensable not only for setting up research studies, but also for storing, organizing and managing enormous quantities of sequencing data. Despite its great advantages, the use of bioinformatics is hindered by difficulties in learning how to use its software tools. The GenEnv database was developed to provide operators involved in biological risk management with a user-friendly tool for sequence analysis. Presently, there are over 20.000 sequence records, and over 9000 bacterial species represented in the database. The initial gene set comprises rDNA16S, rpoB, gyrB. The system allows sequence-driven microbe identification as well as the development of study protocols for research on specific microbe species. Nucleotide sequences are represented graphically. The GenEnv database was designed as a tool for public health operators but also offers wide prospects for scientific research.

Computational Biology↗

A new network-based biologic database system.

Bioinformatics is playing an increasingly important role in the processing and analysis of biomedical data. The collection, storage and analysis of biologic information are key components of bioinformatics. In the genome era, with the explosion of sequence and structural information available to researchers, the structure of biologic databases is becoming more complicated, the contents is becoming larger, and the management and development is also becoming more difficult. A network-based biologic database has been developed. This database system integrates administration, development and analysis of bioinformatics with common and friendly interface. For the life scientists without much computer programming expertise, the system is easy to master.

Computational Biology↗

Novel retinal genes discovered by mining the mouse embryonic RetinalExpress database.

PURPOSE: Bioinformatics has emerged as a powerful tool for identifying novel genes and pathways associated with retinal biology and disease. The developing mouse retina expresses an exceedingly large and complex variety of genes. Many of these genes have not been characterized but nevertheless are likely to have important developmental or physiological functions. The purpose of this study was to use an in silico approach with a mouse embryonic retinal database of cDNAs/expressed sequence tags (ESTs) named RetinalExpress to identify previously uncharacterized genes that are represented in the developing retina. METHODS: cDNA clones unique to the RetinalExpress database were identified by comparing clones in the RetinalExpress database with those in other cDNA/EST databases. We used a hierarchical filtering procedure with high stringency criteria that included sequence quality, colinearity with hypothetical gene sequences, and absence of any substantial existing annotation to select clones that were likely to represent novel genes. Selected clones were located on mouse chromosomes using National Center for Biotechnology Informatics Map Viewer software and the database from the University of California at Santa Cruz Genome Bioinformatics Web browser. The expression of selected retinal transcripts was determined using reverse transcriptase (RT)-PCR. In situ hybridization of sectioned embryonic and postnatal retinas was performed to determine spatial expression patterns of selected transcripts. RESULTS: Of the 27,765 cDNA clones from RetinalExpress that we filtered through several public cDNA/EST databases, 26 cDNA/EST sequences were identified that, at the time of the analysis, were unique to RetinalExpress. Seventeen clones were selected for RT-PCR analysis, and retinal transcripts corresponding to previously uncharacterized genes were unambiguously detected for six clones. Three genes encoded open reading frames containing putative functional domains; one sequence contained an HMG DNA binding domain, another, an RFX DNA binding domain, and another, a phospholipase C catalytic domain X. Transcripts from the genes encoding DNA binding domains were expressed in embryonic and postnatal retinas with distinct spatial patterns. CONCLUSIONS: The characterization of 26 mouse genes whose partial nucleotide sequences were uniquely represented in the RetinalExpress cDNA/EST database demonstrated the feasibility of retinal gene discovery using in silico analysis. Two of these genes had distinctive spatial expression patterns in the retina and one was likely to function as a DNA binding protein in embryonic and postnatal retinas. The gene identification approach described here demonstrates the usefulness of establishing large cDNA/EST databases from highly specialized neuronal tissues such as the retina to find novel genes.

Animals↗

Database-assisted promoter analysis.

The analysis of regulatory sequences is greatly facilitated by database-assisted bioinformatic approaches. The TRANSFAC database contains information on transcription factors and their origins, functional properties and sequence-specific binding activities. Software tools enable us to screen the database with a given DNA sequence for interacting transcription factors. If a regulatory function is already attributed to this sequence then the database-assisted identification of binding sites for proteins or protein classes and subsequent experimental verification might establish functionally relevant sites within this sequence. The binding transcription factors and interacting factors might already be present in the database.

Binding Sites↗

Functional bioinformatics: the cellular response database.

Biological Scientists function in an increasingly data rich environment. The emerging field of bioinformatics is attempting to insure that this flow of information can be structured to support the generation of significant biological hypothesis and ultimately new knowledge. To date, most of the current databases have focused on protein and nucleic acid sequence information as the principle type of data stored for further interpretation. In this paper, we describe the Cellular Response Database. This database stores functional information regarding the changes of cellular gene expression associated with various stimuli, and supports queries linking cell types, expressed genes, and inducers. The database is designed to support information-intensive queries to aid in the determination of biological function, and is flexible enough to allow the storage of a broad range of experimental data such as cytotoxicity data, immunoassays of target gene protein expression, and others.

Cell Physiological Phenomena↗

Linking tumor cell cytotoxicity to mechanism of drug action: an integrated analysis of gene expression, small-molecule screening and structural databases.

An integrated, bioinformatic analysis of three databases comprising tumor-cell-based small molecule screening data, gene expression measurements, and PDB (Protein Data Bank) ligand-target structures has been developed for probing mechanism of drug action (MOA). Clustering analysis of GI50 profiles for the NCI's database of compounds screened across a panel of tumor cells (NCI60) was used to select a subset of unique cytotoxic responses for about 4000 small molecules. Drug-gene-PDB relationships for this test set were examined by correlative analysis of cytotoxic response and differential gene expression profiles within the NCI60 and structural comparisons with known ligand-target crystallographic complexes. A survey of molecular features within these compounds finds thirteen conserved Compound Classes, each class exhibiting chemical features important for interactions with a variety of biological targets. Protein targets for an additional twelve Compound Classes could be directly assigned using drug-protein interactions observed in the crystallographic database. Results from the analysis of constitutive gene expressions established a clear connection between chemo-resistance and overexpression of gene families associated with the extracellular matrix, cytoskeletal organization, and xenobiotic metabolism. Conversely, chemo-sensitivity implicated overexpression of gene families involved in homeostatic functions of nucleic acid repair, aryl hydrocarbon metabolism, heat shock response, proteasome degradation and apoptosis. Correlations between chemo-responsiveness and differential gene expressions identified chemotypes with nonselective (i.e., many) molecular targets from those likely to have selective (i.e., few) molecular targets. Applications of data mining strategies that jointly utilize tumor cell screening, genomic, and structural data are presented for hypotheses generation and identifying novel anticancer candidates.

Antineoplastic Agents↗

Bioinformatics.

Computer databases, networks and software tools are essential materials and methods for biomedical research and are involved in almost every aspect of disease gene mapping and positional cloning. Public databases of DNA and protein sequences and genetic and physical map information are increasing rapidly in size and complexity and are also improving in quality, comprehensiveness, interoperability and access. A new generation of software tools for navigating through the biomedical literature has become available. Programs for sequence homology searching and genetic map construction have become more sophisticated, yet easier to use. Global computer networks are bringing state-of-the-art capabilities to all.

Chromosome Mapping↗