Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “bioinformatic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

ChimerDB--a knowledgebase for fusion sequences.

Chromosome translocation and gene fusion are frequent events in the human genome and are often the cause of many types of tumor. ChimerDB is the database of fusion sequences encompassing bioinformatics analysis of mRNA and expressed sequence tag (EST) sequences in the GenBank, manual collection of literature data and integration with other known database such as OMIM. Our bioinformatics analysis identifies the fusion transcripts that have non-overlapping alignments at multiple genomic loci. Fusion events at exon-exon borders are selected to filter out the cloning artifacts in cDNA library preparation. The result is classified into two groups--genuine chromosome translocation and fusion between neighboring genes owing to intergenic splicing. We also integrated manually collected literature and OMIM data for chromosome translocation as an aid to assess the validity of each fusion event. The database is available at http://genome.ewha.ac.kr/ChimerDB/ for human, mouse and rat genomes.

Animals↗

The Integr8 project--a resource for genomic and proteomic data.

Integr8 (http://www.ebi.ac.uk/integr8/) is providing an integration layer for the exploitation of genomic and proteomic data by drawing on databases maintained at major bioinformatics centres in Europe. Main aims are to store the relationships of biological entities to each other and to entries in other databases, to provide a framework that allows for new kinds of data to be integrated, and to offer an entity-centric view of complete genomes and proteomes. Basic tools for data integration comprise the Proteome Analysis database, the International Protein Index (IPI), the Universal Protein sequence archive (UniParc) and the Genome Reviews. Entry points for the Integr8 portal depend on the users entity of interest: from browsing the taxonomy or with a predetermined species of interest, the species page can be used, and a simple search page leads to different applications when looking for certain protein sequences or genes. Customisable statistics data are available from the BioMart application, and pre-prepared data can be downloaded from the FTP site.

Computational Biology↗

Bioinformatic analysis of autism positional candidate genes using biological databases and computational gene network prediction.

Common genetic disorders are believed to arise from the combined effects of multiple inherited genetic variants acting in concert with environmental factors, such that any given DNA sequence variant may have only a marginal effect on disease outcome. As a consequence, the correlation between disease status and any given DNA marker allele in a genomewide linkage study tends to be relatively weak and the implicated regions typically encompass hundreds of positional candidate genes. Therefore, new strategies are needed to parse relatively large sets of 'positional' candidate genes in search of actual disease-related gene variants. Here we use biological databases to identify 383 positional candidate genes predicted by genomewide genetic linkage analysis of a large set of families, each with two or more members diagnosed with autism, or autism spectrum disorder (ASD). Next, we seek to identify a subset of biologically meaningful, high priority candidates. The strategy is to select autism candidate genes based on prior genetic evidence from the allelic association literature to query the known transcripts within the 1-LOD (logarithm of the odds) support interval for each region. We use recently developed bioinformatic programs that automatically search the biological literature to predict pathways of interacting genes (PATHWAYASSIST and GENEWAYS). To identify gene regulatory networks, we search for coexpression between candidate genes and positional candidates. The studies are intended both to inform studies of autism, and to illustrate and explore the increasing potential of bioinformatic approaches as a compliment to linkage analysis.

Autistic Disorder↗

Linking genomics to immunotherapy by reverse immunology--'immunomics' in the new millennium.

The disclosure of the human genome sequence and rapid advances in genomic expression profiling have revolutionized our knowledge about molecular changes in malignant diseases. Rapidly growing gene expression databases and improvements in bioinformatics tools set the stage for new approaches using large-scale molecular information to develop specific therapeutics in cancer. On one hand, the ability to detect clusters of genes differentially expressed in normal and malignant tissue may lead to widely applicable targeting of defined molecular structures. On the other hand, analyzing the 'molecular fingerprint' of an individual tumor raises the possibility of developing customized therapeutics. One approach to use the emerging new datasets for the development of novel therapeutics is to identify genes that are specifically expressed in tumors as targets for immune intervention. This review will focus on the process from in silico analysis of expression databases and screening of potential candidate genes by bioinformatics to the in vitro and in vivo analysis to determine the immunogenicity of candidate tumor antigens. Basic biological principles of 'reverse immunology' as well as technical advantages and difficulties will be addressed.

Algorithms↗

PatGen--a consolidated resource for searching genetic patent sequences.

UNLABELLED: Compared to the wealth of online resources covering genomic, proteomic and derived data the Bioinformatics community is rather underserved when it comes to patent information related to biological sequences. The current online resources are either incomplete or rather expensive. This paper describes, PatGen, an integrated database containing data from bioinformatic and patent resources. This effort addresses the inconsistency of publicly available genetic patent data coverage by providing access to a consolidated dataset. AVAILABILITY: PatGen can be searched at http://www.patgendb.com CONTACT: rjdrouse@patentinformatics.com.

Abstracting and Indexing↗

ISYS: a decentralized, component-based approach to the integration of heterogeneous bioinformatics resources.

MOTIVATION: Heterogeneity of databases and software resources continues to hamper the integration of biological information. Top-down solutions are not feasible for the full-scale problem of integration across biological species and data types. Bottom-up solutions so far have not integrated, in a maximally flexible way, dynamic and interactive graphical-user-interface components with data repositories and analysis tools. RESULTS: We present a component-based approach that relies on a generalized platform for component integration. The platform enables independently-developed components to synchronize their behavior and exchange services, without direct knowledge of one another. An interface-based data model allows the exchange of information with minimal component interdependency. From these interactions an integrated system results, which we call ISYSf1.gif" BORDER="0">. By allowing services to be discovered dynamically based on selected objects, ISYS encourages a kind of exploratory navigation that we believe to be well-suited for applications in genomic research.

Arabidopsis↗

A compilation of molecular biology web servers: 2006 update on the Bioinformatics Links Directory.

The Bioinformatics Links Directory is a public online resource that lists the servers published in this and all previously published Nucleic Acids Research Web Server issues together with other useful tools, databases and resources for bioinformatics and molecular biology research. This rich directory of tools and websites can be browsed and searched with all listed links freely accessible to the public. The 2006 update includes the 149 websites highlighted in the July 2006 issue of Nucleic Acids Research and brings the total number of servers listed in the Bioinformatics Links Directory to over 1000 links. To aid navigation through this growing resource, all link entries contain a brief synopsis, a citation list and are classified by function in descriptive biological categories. The most up-to-date version of this actively maintained listing of bioinformatics resources is available at the Bioinformatics Links Directory website, http://bioinformatics.ubc.ca/resources/links_directory/. A complete list of all links listed in this Nucleic Acids Research 2006 Web Server issue can be accessed online at http://bioinformatics.ubc.ca/resources/links_directory/narweb2006/. The 2006 update of the Bioinformatics Links Directory, which includes the Web Server list and summaries, is also available online at the Nucleic Acids Research website, http://nar.oupjournals.org/.

Computational Biology↗

E-MSD: improving data deposition and structure quality.

The Macromolecular Structure Database (MSD) (http://www.ebi.ac.uk/msd/) [H. Boutselakis, D. Dimitropoulos, J. Fillon, A. Golovin, K. Henrick, A. Hussain, J. Ionides, M. John, P. A. Keller, E. Krissinel et al. (2003) E-MSD: the European Bioinformatics Institute Macromolecular Structure Database. Nucleic Acids Res., 31, 458-462.] group is one of the three partners in the worldwide Protein DataBank (wwPDB), the consortium entrusted with the collation, maintenance and distribution of the global repository of macromolecular structure data [H. Berman, K. Henrick and H. Nakamura (2003) Announcing the worldwide Protein Data Bank. Nature Struct. Biol., 10, 980.]. Since its inception, the MSD group has worked with partners around the world to improve the quality of PDB data, through a clean up programme that addresses inconsistencies and inaccuracies in the legacy archive. The improvements in data quality in the legacy archive have been achieved largely through the creation of a unified data archive, in the form of a relational database that stores all of the data in the wwPDB. The three partners are working towards improving the tools and methods for the deposition of new data by the community at large. The implementation of the MSD database, together with the parallel development of improved tools and methodologies for data harvesting, validation and archival, has lead to significant improvements in the quality of data that enters the archive. Through this and related projects in the NMR and EM realms the MSD continues to improve the quality of publicly available structural data.

Computational Biology↗

Psychiatric genetics in silico: databases and tools for psychiatric geneticists.

Bioinformatics can significantly impact the laboratory genetics process from the study design phase to conclusive identification of a disease gene. The present review will highlight key databases to enhance psychiatric genetic study design, based on full use of genomics data and the golden path sequence. It will address methods to ensure comprehensive genetic data mining, using the best available genomic and genetic databases such as the University of California Santa Cruz human genome browser, Ensembl, Mapview, dbSNP and GDB, and locus-specific databases such as Online Mendelian Inheritance In Man. Using the golden path sequence as a template, with the necessary quality checks, it is possible to design detailed genetic studies from sequence information alone. Drawing together this diverse information, it is possible to characterize a locus or gene in silico to a very detailed level. This in turn can have real cost and efficiency benefits by assisting in the identification of markers that are most likely to be informative, or by highlighting the best candidate genes for study.

Databases, Bibliographic↗

T-cell epitopes within the complementarity-determining and framework regions of the tumor-derived immunoglobulin heavy chain in multiple myeloma.

The idiotypic structure of the monoclonal immunoglobulin (Ig) in multiple myeloma (MM) might be regarded as a tumor-specific antigen. The present study was designed to identify T-cell epitopes of the variable region of the Ig heavy chain (VH) in MM (n = 5) using bioinformatics and analyze the presence of naturally occurring T cells against idiotype-derived peptides. A large number of human-leukocyte-antigen (HLA)-binding (class I and II) peptides were identified. The frequency of predicted epitopes depended on the database used: 245 in bioinformatics and molecular analysis section (BIMAS) and 601 in SYFPEITHI. Most of the peptides displayed a binding half-life or score in the low or intermediate affinity range. The majority of the predicted peptides were complementarity-determining region (CDR)-rather than framework region (FR)-derived (52%-60% vs 40%-48%, respectively). Most of the predicted peptides were confined to the CDR2-FR3-CDR3 "geographic" region of the Ig-VH region (70%), and significantly fewer peptides were found within the flanking (FR1-CDR1-FR2 and FR4) regions (P <.01). There were 8- to 10-amino acid (aa) long peptides corresponding to the CDRs and fitting to the actual HLA-A/B haplotypes that spontaneously recognized, albeit with a low magnitude, type I T cells (interferon gamma), indicating an ongoing major histocompatibility complex (MHC) class I-restricted T-cell response. Most of those peptides had a low binding half-life (BIMAS) and a low/intermediate score (SYFPEITHI). Furthermore, 15- to 20-aa long CDR1-3-derived peptides also spontaneously recognized type I T cells, indicating the presence of MHC class II-restricted T cells as well. This study demonstrates that a large number of HLA-binding idiotypic peptides can be identified in patients with MM. Such peptides may spontaneously induce a type I MHC class I- as well as class II-restricted memory T-cell response.

Aged↗

Integrated Bioinformatics Analysis Revealing that the NSDHL Gene Might Be Associated with the Progression of Western HFD/SW-Induced Hepatocellular Carcinoma.

BACKGROUND AND OBJECTIVE: Hepatocellular carcinoma (HCC) remains a significant global health concern. However, the etiology and pathogenesis of HCC have yet to be fully elucidated. Previous studies have indicated a close association between obesity and the occurrence and progression of HCC. The objective of this study was to employ bioinformatics strategies in order to explore key genes associated with the clinical diagnosis and prognosis of HCC induced by a Western high-fat diet and sugar water (HFD/SW). MATERIALS AND METHODS: We obtained the expression profile chip data GSE197884 from the Gene Expression Omnibus (GEO) database. Subsequently, &#x201c;DESeq&#x201d; and &#x201c;Limma&#x201d; R packages were employed to identify differentially expressed genes (DEGs) while constructing a co-expressed gene network using weighted gene co-expression analysis (WGCNA). Functional enrichment analyses were then carried out, followed by the construction of a protein-protein interaction (PPI) network to uncover core genes. The core genes were confirmed through data retrieved from The Cancer Genome Atlas (TCGA) database in order to determine their status as hub genes. Finally, survival and tumor immune infiltration analyses were performed to unveil the prognostic significance of these hub genes. RESULTS: In total, 126 intersection targets were retrieved through the Venn diagram. Gene ontology (GO) enrichment and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analyses revealed that the DEGs were primarily related to the proliferation and apoptosis of HCC cells, the digestion and metabolism of liver cells, the HCC tumor microenvironment, and immune response. The PPI network analysis identified 11 core targets, among which seven hub genes, including NSDHL, MVK, SQLW, GCAT, ALAS2, GLDC, and AGXT, were obtained after TCGA database validation. Furthermore, it was found that NSDHL was closely associated with the clinical diagnosis and prognosis of HCC induced by HFD/SW and also affected the cellular immune infiltration in the HCC tumor microenvironment. CONCLUSION: The present study demonstrated a significantly elevated expression of NSDHL in HCC tissues, suggesting its potential as a specific biomarker for precise clinical diagnosis and prognosis assessment of HCC induced by HFD/SW.

Computational Biology↗

DBCollHIV: a database system for collaborative HIV analysis in Brazil.

We developed a database system for collaborative HIV analysis (DBCollHIV) in Brazil. The main purpose of our DBCollHIV project was to develop an HIV-integrated database system with analytical bioinformatics tools that would support the needs of Brazilian research groups for data storage and sequence analysis. Whenever authorized by the principal investigator, this system also allows the integration of data from different studies and/or the release of the data to the general public. The development of a database that combines sequences associated with clinical/epidemiological data is difficult without the active support of interdisciplinary investigators. A functional database that securely stores data and helps the investigator to manipulate their sequences before publication would be an attractive tool for investigators depositing their data and collaborating with other groups. DBCollHIV allows investigators to manipulate their own datasets, as well as integrating molecular and clinical HIV data, in an innovative fashion.

Brazil↗

Genotypic drug resistance interpretation systems--the cutting edge of antiretroviral therapy.

The technical quality of genotypic and phenotypic drug resistance testing has considerably improved, and therefore the major challenge now lies in the interpretation of drug resistance. This is due to several facts: (i) in times of combination therapy, the effect of drug resistance-associated mutations cannot be considered independently, (ii) many additive and subtractive interactions between mutations exist, and resistant strains may exhibit varying degrees of cross-resistance, (iii) the phenotype cannot adequately determine slight, but clinically relevant, differences for those drugs with a narrow range of resistance, and (iv) pharmacokinetic interactions may shift relevant levels of drug resistance. Genotypic drug resistance interpretation systems are designed to solve these problems. Rule-based systems incorporate current knowledge about correlations between genotype, phenotype and clinical response. Database-driven systems use the information provided by paired geno- and phenotypic data, applying database matching search or bioinformatic approaches. For detailed comparison, 11 interpretation systems were selected which present a comprehensive system for most of the available drugs, can easily be accessed via the Internet and are regularly updated. The systems were characterized for the source data, access, input, output, and availability of clinical studies. For further comparison, existing clinical databases should be merged into one large database to allow competition between the systems. This may also solve the burning problem of clinically relevant cut-offs. Head-to-head comparisons of interpretation systems require large prospective randomized trials in which only the interpretation system is different between groups, before a consensus can be achieved for the best antiretroviral therapy of the individual patient.

Algorithms↗

Developmental bioinformatics: linking genetic data to virtual embryos.

This paper discusses current efforts to produce databases of gene expression for the major model embryos used in developmental biology. The efforts to build these resources were motivated by the need for immediate internet access to all types of research data, and the production of these databases is a major and new challenge for bioinformatics. Thus far bioinformatics has mainly been concerned with textually oriented resources and data, much of it concerned with gene and protein sequences. Because the genetic basis of developmental biology is integrated with developmental anatomy, these databases require the use of images to link molecular data with spatial information. In order to standardise database formats, digital atlases of some model systems are being produced that include integrated anatomical descriptions and these are being linked to appropriate genetic data. Integrating such image-based, searchable data into databases makes new demands on the field of bioinformatics and we consider here the imaging modalities that are used to obtain information and we discuss in particular the production of 3D images from serial sections. Next, we consider how to integrate textual and spatial descriptions of gene expression and the key tool needed to make this possible, i.e. anatomical nomenclature. A short review of internet resources on developmental biology is also given and future prospects for the development of these databases are discussed.

Animals↗

Mutational data integration in gene-oriented files of the Hermansky-Pudlak Syndrome database.

Hermansky-Pudlak Syndrome (HPS) is a genetically heterogeneous disorder characterized by oculocutaneous albinism and prolonged bleeding due to abnormal vesicle trafficking to lysosomes and related organelles such as melanosomes and platelet dense granules. This HPS database (HPSD; http://liweilab.genetics.ac.cn/HPSD/) provides integrated, annotatory, and curative data that is distributed in a variety of public databases or predicted by bioinformatics servers for the recently cloned human and mouse HPS genes, as well as for the genes responsible for HPSrelated syndromes, such as ChediakHigashi Syndrome (CHS), Griscelli syndrome (GS), oculocutaneous albinism (OCA), Usher syndrome type 1B (USH1B), and ocular albinism (OA). The HPSD is designed by using a unique GeneOriented File (GOF) format. Seven blocks (genomic, transcript, protein, function, mutation, phenotype, and reference) are carefully annotated in each userfriendly GOF entry. The HPSD emphasizes paired human and mouse GOF entries. The genes included in this database (currently 58 in total) are arbitrarily divided into four categories: 1) Human and Mouse HPS, 2) Mouse HPS Only, 3) Putative Mouse or Human HPS, and 4) HPS Related Syndromes. All the mutations in these genes are integrated in the GOFs. We expect that these very informative and peerreviewed GOFs will be shortcuts to utilize the webbased information for the emerging interdisciplinary studies of HPS.

Animals↗

The Androgen Receptor Gene Mutations Database.

The current version of the androgen receptor (AR) gene mutations database is described. The total number of reported mutations has risen from 272 to 309 in the past year. We have expanded the database: (i) by giving each entry an accession number; (ii) by adding information on the length of polymorphic polyglutamine (polyGln) and polyglycine (polyGly) tracts in exon 1; (iii) by adding information on large gene deletions; (iv) by providing a direct link with a completely searchable database (courtesy EMBL-European Bioinformatics Institute). The addition of the exon 1 polymorphisms is discussed in light of their possible relevance as markers for predisposition to prostate or breast cancer. The database is also available on the internet (http://www.mcgill. ca/androgendb/ ), from EMBL-European Bioinformatics Institute (ftp. ebi.ac.uk/pub/databases/androgen ), or as a Macintosh FilemakerPro or Word file (MC33@musica.mcgill.ca).

Computer Communication Networks↗

Comprehensive gene expression profile of the adult human renal cortex: analysis by cDNA array hybridization.

BACKGROUND: Profiling of gene expression in healthy and diseased renal tissue is important for elucidating the pathogenesis of renal diseases. Comprehensive information about the genes expressed in renal tissue is unavailable. The recently developed cDNA array hybridization methodology allows simultaneous monitoring of thousands of genes expressed renal tissue. METHODS: Complex [alpha-33P]-labeled cDNA probes were prepared from histopathologically uninvolved remnants of nine renal tissues obtained by nephrectomy. Each probe was hybridized to a high-density array of 18,326 paired target genes. The radioactive hybridization signals by phosphorimager screens were quantitated by special software. Bioinformatics from public genomic databases were used to assign a chromosomal location of each expressed transcript and gene function. Cluster analysis was used to arrange genes according to the similarity in pattern of gene expression. RESULTS: A total of 7563 different gene transcripts was detected in the nine tissue samples. Approximately 870 of these genes were full-length mRNA human transcripts (HT), and the remaining 6693 were expressed sequence tags (ESTs). The full-length transcripts were classified by function of the gene product and were listed with information of their chromosomal positions. To allow a comparison between gene expression in clinical and experimental studies, the mouse genes with known similar function to the human counterpart were included in the bioinformatics analysis. Cluster analysis of 502 full-length genes that are expressed in four or more renal tissues revealed more than 110 genes that are highly expressed in all the renal specimens. CONCLUSIONS: The presented data constitute a comprehensive preliminary transcriptional map of the adult human renal cortex. The information may serve as a resource for speeding up the discovery of genes underlying human renal disease. The integrated listing of the full-length expressed human and mouse genes is available through e-mail (Abdalla_Rifai@Brown.edu).

Adult↗

wFleaBase: the Daphnia genome database.

BACKGROUND: wFleaBase is a database with the necessary infrastructure to curate, archive and share genetic, molecular and functional genomic data and protocols for an emerging model organism, the microcrustacean Daphnia. Commonly known as the water-flea, Daphnia's ecological merit is unequaled among metazoans, largely because of its sentinel role within freshwater ecosystems and over 200 years of biological investigations. By consequence, the Daphnia Genomics Consortium (DGC) has launched an interdisciplinary research program to create the resources needed to study genes that affect ecological and evolutionary success in natural environments. DISCUSSION: These tools include the genome database wFleaBase, which currently contains functions to search and extract information from expressed sequenced tags, genome survey sequences and full genome sequencing projects. This new database is built primarily from core components of the Generic Model Organism Database project, and related bioinformatics tools. SUMMARY: Over the coming year, preliminary genetic maps and the nearly complete genomic sequence of Daphnia pulex will be integrated into wFleaBase, including gene predictions and ortholog assignments based on sequence similarities with eukaryote genes of known function. wFleaBase aims to serve a large ecological and evolutionary research community. Our challenge is to rapidly expand its content and to ultimately integrate genetic and functional genomic information with population-level responses to environmental challenges. URL: http://wfleabase.org/.

Animals↗