Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “bioinformatic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

The Radiation Hybrid Database.

Since July 1995, the European Bioinformatics Institute (EBI) has maintained RHdb, a public database for radiation hybrid data. Radiation hybrid data are used in the generation of alternative genetic maps as they can include non-polymorphic markers and are also powerful enough to order unresolved genetic clusters of polymorphic STSs. The EBI is an Outstation of the European Molecular Biology Laboratory (EMBL).

Academies and Institutes↗

The radiation hybrid database.

Since July 1995, the European Bioinformatics Institute (EBI) has maintained the Radiation Hybrid database (RHdb; http://www.ebi.ac. uk/RHdb ), a public database for radiation hybrid data. Radiation hybrid mapping is an important technique for determining high resolution maps. Recently, CORBA access has been added to RHdb. The EBI is an Outstation of the European Molecular Biology Laboratory (EMBL).

Animals↗

Bioinformatics in microbial biotechnology--a mini review.

The revolutionary growth in the computation speed and memory storage capability has fueled a new era in the analysis of biological data. Hundreds of microbial genomes and many eukaryotic genomes including a cleaner draft of human genome have been sequenced raising the expectation of better control of microorganisms. The goals are as lofty as the development of rational drugs and antimicrobial agents, development of new enhanced bacterial strains for bioremediation and pollution control, development of better and easy to administer vaccines, the development of protein biomarkers for various bacterial diseases, and better understanding of host-bacteria interaction to prevent bacterial infections. In the last decade the development of many new bioinformatics techniques and integrated databases has facilitated the realization of these goals. Current research in bioinformatics can be classified into: (i) genomics--sequencing and comparative study of genomes to identify gene and genome functionality, (ii) proteomics--identification and characterization of protein related properties and reconstruction of metabolic and regulatory pathways, (iii) cell visualization and simulation to study and model cell behavior, and (iv) application to the development of drugs and anti-microbial agents. In this article, we will focus on the techniques and their limitations in genomics and proteomics. Bioinformatics research can be classified under three major approaches: (1) analysis based upon the available experimental wet-lab data, (2) the use of mathematical modeling to derive new information, and (3) an integrated approach that integrates search techniques with mathematical modeling. The major impact of bioinformatics research has been to automate the genome sequencing, automated development of integrated genomics and proteomics databases, automated genome comparisons to identify the genome function, automated derivation of metabolic pathways, gene expression analysis to derive regulatory pathways, the development of statistical techniques, clustering techniques and data mining techniques to derive protein-protein and protein-DNA interactions, and modeling of 3D structure of proteins and 3D docking between proteins and biochemicals for rational drug design, difference analysis between pathogenic and non-pathogenic strains to identify candidate genes for vaccines and anti-microbial agents, and the whole genome comparison to understand the microbial evolution. The development of bioinformatics techniques has enhanced the pace of biological discovery by automated analysis of large number of microbial genomes. We are on the verge of using all this knowledge to understand cellular mechanisms at the systemic level. The developed bioinformatics techniques have potential to facilitate (i) the discovery of causes of diseases, (ii) vaccine and rational drug design, and (iii) improved cost effective agents for bioremediation by pruning out the dead ends. Despite the fast paced global effort, the current analysis is limited by the lack of available gene-functionality from the wet-lab data, the lack of computer algorithms to explore vast amount of data with unknown functionality, limited availability of protein-protein and protein-DNA interactions, and the lack of knowledge of temporal and transient behavior of genes and pathways.

Journal Article↗

Bioinformatics analyses of circular dichroism protein reference databases.

MOTIVATION: Circular dichroism (CD) spectroscopy has become established as a key method for determining the secondary structure contents of proteins which has had a significant impact on molecular biology. Many excellent mathematical protocols have been developed for this purpose and their quality is above question. However, reference database sets of proteins, with CD spectra matched to secondary structure components derived from X-ray structures, provide the key resource for this task. These databases were created many years ago, before most CD spectrophotometers became standardized and before it was commonplace to validate X-ray structures prior to publication. The analyses presented here were undertaken to investigate the overall quality of these reference databases in light of their extensive usage in determining protein secondary structure content from CD spectra. RESULTS: The analyses show that there are a number of significant problems associated with the CD reference database sets in current use. There are disparities between CD spectra for the same protein collected by different groups. These include differences in magnitudes, peak positions or both. However, many current reference sets are now amalgamations of spectra from these groups, introducing inconsistencies that can lead to inaccuracies in the determination of secondary structure components from the CD spectra. A number of the X-ray structures used fall short on the validation criteria now employed as standard for structure determination. Many have substantial percentages of residues in the disallowed regions of the Ramachandran plot. Hence their calculated secondary structure components, used as a foundation for the reference databases, are likely to be in error. Additionally, the coverage of secondary structure space in the reference datasets is poorly correlated to the secondary structure components found in the Protein Data Bank. A conclusion is that a new reference CD database with cross-correlated, machine-independent CD spectra and validated X-ray structures that cover more secondary structure components, including diverse protein folds, is now needed. However, that reasonably accurate values for the secondary structure content of proteins can be determined from spectra is a testament to CD spectroscopy being a very powerful technique.

Circular Dichroism↗

The Kidney Development Database.

The Kidney Development Database is a bioinformatics resource dedicated to providing easily accessible information on gene expression during the development of the pro-, meso-, and metanephroi of a range of vertebrates. It also contains data on mutant phenotypes and on the effects of experimental manipulation of kidneys developing in culture. The database is searchable by gene name or by expression pattern. It is now being used as a test bed for more "advanced" search strategies that measure hypotheses of gene interactions against expression data to test for any clashes that would make the hypotheses untenable and that scan the database for potentially interesting correlations between changes in gene expression. The Kidney Development Database can be accessed free of charge via the World Wide Web at either of the following uniform resource locators (URLs); http://golgi.ana.ed.ac.uk/kidhome.html, and http://www.ana.ed.ac.uk/anatomy/database/kidbase/ kidhome.html.

Animals↗

A complementary bioinformatics approach to identify potential plant cell wall glycosyltransferase-encoding genes.

Plant cell wall (CW) synthesizing enzymes can be divided into the glycan (i.e. cellulose and callose) synthases, which are multimembrane spanning proteins located at the plasma membrane, and the glycosyltransferases (GTs), which are Golgi localized single membrane spanning proteins, believed to participate in the synthesis of hemicellulose, pectin, mannans, and various glycoproteins. At the Carbohydrate-Active enZYmes (CAZy) database where e.g. glucoside hydrolases and GTs are classified into gene families primarily based on amino acid sequence similarities, 415 Arabidopsis GTs have been classified. Although much is known with regard to composition and fine structures of the plant CW, only a handful of CW biosynthetic GT genes-all classified in the CAZy system-have been characterized. In an effort to identify CW GTs that have not yet been classified in the CAZy database, a simple bioinformatics approach was adopted. First, the entire Arabidopsis proteome was run through the Transmembrane Hidden Markov Model 2.0 server and proteins containing one or, more rarely, two transmembrane domains within the N-terminal 150 amino acids were collected. Second, these sequences were submitted to the SUPERFAMILY prediction server, and sequences that were predicted to belong to the superfamilies NDP-sugartransferase, UDP-glycosyltransferase/glucogen-phosphorylase, carbohydrate-binding domain, Gal-binding domain, or Rossman fold were collected, yielding a total of 191 sequences. Fifty-two accessions already classified in CAZy were discarded. The resulting 139 sequences were then analyzed using the Three-Dimensional-Position-Specific Scoring Matrix and mGenTHREADER servers, and 27 sequences with similarity to either the GT-A or the GT-B fold were obtained. Proof of concept of the present approach has to some extent been provided by our recent demonstration that two members of this pool of 27 non-CAZy-classified putative GTs are xylosyltransferases involved in synthesis of pectin rhamnogalacturonan II (J. Egelund, B.L. Petersen, A. Faik, M.S. Motawia, C.E. Olsen, T. Ishii, H. Clausen, P. Ulvskov, and N. Geshi, unpublished data).

Amino Acid Motifs↗

The ASAP II database: analysis and comparative genomics of alternative splicing in 15 animal species.

We have greatly expanded the Alternative Splicing Annotation Project (ASAP) database: (i) its human alternative splicing data are expanded approximately 3-fold over the previous ASAP database, to nearly 90,000 distinct alternative splicing events; (ii) it now provides genome-wide alternative splicing analyses for 15 vertebrate, insect and other animal species; (iii) it provides comprehensive comparative genomics information for comparing alternative splicing and splice site conservation across 17 aligned genomes, based on UCSC multigenome alignments; (iv) it provides an approximately 2- to 3-fold expansion in detection of tissue-specific alternative splicing events, and of cancer versus normal specific alternative splicing events. We have also constructed a novel database linking orthologous exons and orthologous introns between genomes, based on multigenome alignment of 17 animal species. It can be a valuable resource for studies of gene structure evolution. ASAP II provides a new web interface enabling more detailed exploration of the data, and integrating comparative genomics information with alternative splicing data. We provide a set of tools for advanced data-mining of ASAP II with Pygr (the Python Graph Database Framework for Bioinformatics) including powerful features such as graph query, multigenome alignment query, etc. ASAP II is available at http://www.bioinformatics.ucla.edu/ASAP2.

Alternative Splicing↗

GLAD: a system for developing and deploying large-scale bioinformatics grid.

MOTIVATION: Grid computing is used to solve large-scale bioinformatics problems with gigabytes database by distributing the computation across multiple platforms. Until now in developing bioinformatics grid applications, it is extremely tedious to design and implement the component algorithms and parallelization techniques for different classes of problems, and to access remotely located sequence database files of varying formats across the grid. In this study, we propose a grid programming toolkit, GLAD (Grid Life sciences Applications Developer), which facilitates the development and deployment of bioinformatics applications on a grid. RESULTS: GLAD has been developed using ALiCE (Adaptive scaLable Internet-based Computing Engine), a Java-based grid middleware, which exploits the task-based parallelism. Two bioinformatics benchmark applications, such as distributed sequence comparison and distributed progressive multiple sequence alignment, have been developed using GLAD.

Computational Biology↗

An inquiry into protein structure and genetic disease: introducing undergraduates to bioinformatics in a large introductory course.

This inquiry-based lab is designed around genetic diseases with a focus on protein structure and function. To allow students to work on their own investigatory projects, 10 projects on 10 different proteins were developed. Students are grouped in sections of 20 and work in pairs on each of the projects. To begin their investigation, students are given a cDNA sequence that translates into a human protein with a single mutation. Each case results in a genetic disease that has been studied and recorded in the Online Mendelian Inheritance in Man (OMIM) database. Students use bioinformatics tools to investigate their proteins and form a hypothesis for the effect of the mutation on protein function. They are also asked to predict the impact of the mutation on human physiology and present their findings in the form of an oral report. Over five laboratory sessions, students use tools on the National Center for Biotechnology Information (NCBI) Web site (BLAST, LocusLink, OMIM, GenBank, and PubMed) as well as ExPasy, Protein Data Bank, ClustalW, the Kyoto Encyclopedia of Genes and Genomes (KEGG) database, and the structure-viewing program DeepView. Assessment results showed that students gained an understanding of the Web-based databases and tools and enjoyed the investigatory nature of the lab.

Algorithms↗

Usability of the kink parameters for nucleic acid structure in database.

DNA/RNA molecules with the specific three-dimensional structure, express the specific structural and biological functions. The knowledge integration based on three dimensional structures can be used to explain biochemical observations, to predict biological functions and to design drugs specific to a given complex system. The database cataloging the interaction motifs of nucleic acid moieties has been developed. Skew matrix, a kind of kink parameter, affords the structural description between the adjacent moieties. The proper values in skew matrix have the good properties, i.e., impregnable and flexible presentation. The geometrical parameters about hydrogen bond and base stacking would also provide the tolerant aspect for stereochemical bioinformatics. The tentative database including these parameters with the species and physical properties of the surrounding nucleic acid components and amino acids has been constructed.

Amino Acids↗

Integr8: enhanced inter-operability of European molecular biology databases.

OBJECTIVES: The increasing production of molecular biology data in the post-genomic era, and the proliferation of databases that store it, require the development of an integrative layer in database services to facilitate the synthesis of related information. The solution of this problem is made more difficult by the absence of universal identifiers for biological entities, and the breadth and variety of available data. METHODS: Integr8 was modelled using UML (Universal Modelling Language). Integr8 is being implemented as an n-tier system using a modern object-oriented programming language (Java). An object-relational mapping tool, OJB, is being used to specify the interface between the upper layers and an underlying relational database. RESULTS: The European Bioinformatics Institute is launching the Integr8 project. Integr8 will be an automatically populated database in which we will maintain stable identifiers for biological entities, describe their relationships with each other (in accordance with the central dogma of biology), and store equivalences between identified entities in the source databases. Only core data will be stored in Integr8, with web links to the source databases providing further information. CONCLUSIONS: Integr8 will provide the integrative layer of the next generation of bioinformatics services from the EBI. Web-based interfaces will be developed to offer gene-centric views of the integrated data, presenting (where known) the links between genome, proteome and phenotype.

Computational Biology↗

In silico analysis of the human kallikrein gene 6.

Kallikreins are a family of 15 serine proteases clustered together on the long arm of chromosome 19. Recent reports have linked kallikreins to malignancy. The human kallikrein gene 6 (KLK6) is a newly characterized member of the human kallikrein gene family. Recent work has focused on the possible role of this gene and its protein product as a tumor marker and its involvement in diseases of the central nervous system. In this study, we performed extensive in silico analyses of KLK6 expression from different databases using various bioinformatic tools. These data enabled us to construct and verify the longest transcript for this kallikrein, to identify several polymorphisms among published sequences and to summarize the 21 single-nucleotide polymorphisms of the gene. Our expressed sequence tag (EST) analyses suggest the existence of seven new splice variants of the gene, in addition to the already reported ones. Most of these variants were identified in libraries from cancerous tissues. KLK6 orthologues were identified from three other species with approximately 86% overall homology with rat and mouse orthologues. We also utilized several databases to compare KLK6 gene expression in normal and cancerous tissues. The serial analysis of gene expression and EST expression profiles showed upregulation of the gene in female genital (ovarian and uterine) and gastrointestinal (gastric, colon, esophageal and pancreatic) cancers. Significant downregulation was observed in breast cancers and brain tumors, in relation to their normal counterparts.

Brain Neoplasms↗

Ontological visualization of protein-protein interactions.

BACKGROUND: Cellular processes require the interaction of many proteins across several cellular compartments. Determining the collective network of such interactions is an important aspect of understanding the role and regulation of individual proteins. The Gene Ontology (GO) is used by model organism databases and other bioinformatics resources to provide functional annotation of proteins. The annotation process provides a mechanism to document the binding of one protein with another. We have constructed protein interaction networks for mouse proteins utilizing the information encoded in the GO annotations. The work reported here presents a methodology for integrating and visualizing information on protein-protein interactions. RESULTS: GO annotation at Mouse Genome Informatics (MGI) captures 1318 curated, documented interactions. These include 129 binary interactions and 125 interaction involving three or more gene products. Three networks involve over 30 partners, the largest involving 109 proteins. Several tools are available at MGI to visualize and analyze these data. CONCLUSIONS: Curators at the MGI database annotate protein-protein interaction data from experimental reports from the literature. Integration of these data with the other types of data curated at MGI places protein binding data into the larger context of mouse biology and facilitates the generation of new biological hypotheses based on physical interactions among gene products.

Animals↗

Integrative proteomics and bioinformatics pipelines for PTM profiling.

Post-translational modifications (PTMs) regulate protein function across all life forms and allow plants to respond rapidly to biotic and abiotic stress. Over 450 PTM types have been described across organisms, of which 23-33 have been experimentally confirmed in plants, including phosphorylation, acetylation, methylation, glycosylation, ubiquitination, and sumoylation. These modifications are highly dynamic and often reversible, and frequently act in combination, or "crosstalk," to fine-tune cellular processes. Advances in high-resolution mass spectrometry and large-scale genome sequencing continue to expand the catalogue of known PTM sites, while machine learning and deep learning approaches increasingly support prediction of PTM site localization and function. Unlike broader surveys of plant PTMs, this review focuses specifically on O-phosphorylation and Lys-N(ε)-acetylation, the two best-characterized and most extensively crosstalking PTMs in plants, and integrates four perspectives: the historical development of proteomic and bioinformatics approaches to these modifications; current mass spectrometry-based workflows and enrichment strategies; the bioinformatics tools and databases available for their analysis; and the technical and species-related challenges, particularly in non-model plants, that currently limit their study. We close by outlining priority directions for future research, including multi-omics integration, AI-based prediction, and the translation of PTM knowledge into crop stress resilience and breeding applications.

Protein Processing, Post-Translational↗

Many human genes are transcribed from the antisense promoter of L1 retrotransposon.

Human L1 retrotransposon has two transcription-regulatory regions: an internal or sense promoter driving transcription of the full-length L1, and an antisense promoter (ASP) driving transcription in the opposite direction into adjacent cellular sequences yielding chimeric transcripts. Both promoters are located in the 5'-untranslated region (5'-UTR) of L1. Chimeric transcripts derived from the L1 ASP are highly represented in expressed-sequence tag (EST) databases. Using a bioinformatics approach, we have characterized 10 chimeric ESTs (cESTs) derived from the EST division of GenBank. These cESTs contained 3' regions similar or identical to known cellular mRNA sequences. They were accurately spliced and preferentially expressed in tumor cell lines. Analysis of the hundreds of cESTs suggests that the L1 ASP-driven transcription is a common phenomenon not only for tumor cells but also for normal ones and may involve transcriptional interference or epigenetic control of different cellular genes.

5' Flanking Region↗

Identification of cDNAs associated with late dedifferentiation in adult newt forelimb regeneration.

Epimorphic limb regeneration in the adult newt involves the dedifferentiation of differentiated cells to yield a pluripotent blastemal cell. These mesenchymal-like cells proliferate and subsequently respond to patterning and differentiation cues to form a new limb. Understanding the dedifferentiation process requires the selective identification of dedifferentiating cells within the heterogeneous population of cells in the regenerate. In this study, representational differences analysis was used to produce an enriched population of dedifferentiation-associated cDNA fragments. Fifty-nine unique cDNA fragments were identified, sequenced, and analyzed using bioinformatics tools and databases. Some of these clones demonstrate significant similarity to known genes in other species. Other clones can be linked by homology to pathways previously implicated in the dedifferentiation process. These data will form the basis for further analyses to elucidate the role of candidate genes in the dedifferentiation process during newt forelimb regeneration.

Aging↗

Approaches for analyzing human mutations and nucleotide sequence variation: a report from the Seventh International Mutation Detection meeting, 2003.

The Seventh International Symposium on Mutations in the Human Genome, Mutation Detection 2003, was held during 2-6 July 2003 in Palm Cove near Cairns, Australia. The meeting was organized under the auspices of the Human Genome Organisation (HUGO) as a satellite meeting of the International World Congress of Genetics, held in Melbourne the following week. Meeting participants reported on advances in mutation detection technologies, including advances in high-throughput detection systems for SNP genotyping applicable to the international haplotype mapping project (HapMap); and bioinformatics tools, including databases for handling and processing growing amounts of genome variation data. This meeting report summarizes the presentations and cites related articles from the special issue of Human Mutation (Volume 23#5, May 2004; available online at www.wiley.com/humanmutation).

Base Sequence↗