Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “bioinformatic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Pharmacogenomics and its potential impact on drug and formulation development.

Recent advances in genomic research have provided the basis for new insights into the importance of genetic and genomic markers during the different stages of drug development. A new field of research, pharmacogenomics, which studies the relationship between drug effects and the genome, has emerged. Structural pharmacogenomics maps the complete DNA sequences of whole genomes (genotypes) including individual variations, and functional pharmacogenomics assesses the expression levels of thousands of genes in one single experiment. Together, these two areas of pharmacogenomics have generated massive databases, which have become a challenge for the research field of informatics and have fostered a new branch of research, bioinformatics. If skillfully used, the databases generated by pharmacogenomics together with data mining on the Web promise to improve the drug development process in a variety of areas: identification of drug targets, evaluation of toxicity, classification of diseases, evaluation of formulations, assessment of drug response and treatment, post-marketing applications, and development of personalized medicines.

Animals↗

Molecular biology of pyridine nucleotide and nicotine biosynthesis.

Nicotinamide adenine dinucleotide (NAD) is a ubiquitous coenzyme in oxidation-reduction reactions. Recent animal and fungal studies show that it also plays important roles in transcriptional regulation, longevity, and age-associated diseases. NAD is synthesized de novo from aspartic acid in E. coli or from tryptophan in animals, by way of quinolinic acid. Although the number of biochemical studies on NAD is very limited, a bioinformatic search of genome databases suggests that Arabidopsis (dicots) synthesizes NAD from aspartic acid whereas rice (monocots) may utilize both aspartate and tryptophan as starting amino acids. The salvage pathway recycles the breakdown products of NAD metabolism. In tobacco, an intermediate in the de novo NAD synthetic pathway supplies the pyridine ring moiety of nicotine alkaloids. Gene expression studies in tobacco suggest that part of the NAD pathway is coordinately regulated with nicotine biosynthesis.

Animals↗

Challenges of DNA profiling in mass disaster investigations.

In cases of mass disaster, there is often a need for managing, analyzing, and comparing large numbers of biological samples and DNA profiles. This requires the use of laboratory information management systems for large-scale sample logging and tracking, coupled with bioinformatic tools for DNA database searching according to different matching algorithms, and for the evaluation of the significance of each match by likelihood ratio calculations. There are many different interrelated factors and circumstances involved in each specific mass disaster scenario that may challenge the final DNA identification goal, such as: the number of victims, the mechanisms of body destruction, the extent of body fragmentation, the rate of DNA degradation, the body accessibility for sample collection, or the type of DNA reference samples availability. In this paper, we examine the different steps of the DNA identification analysis (DNA sampling, DNA analysis and technology, DNA database searching, and concordance and kinship analysis) reviewing the "lessons learned" and the scientific progress made in some mass disaster cases described in the scientific literature. We will put special emphasis on the valuable scientific feedback that genetic forensic community has received from the collaborative efforts of several public and private USA forensic laboratories in assisting with the more critical areas of the World Trade Center (WTC) mass fatality of September 11, 2001. The main challenges in identifying the victims of the recent South Asian Tsunami disaster, which has produced the steepest death count rise in history, will also be considered. We also present data from two recent mass fatality cases that involved Spanish victims: the Madrid terrorist attack of March 11, 2004, and the Yakolev-42 aircraft accident in Trabzon, Turkey, of May 26, 2003.

DNA Fingerprinting↗

EBI databases and services.

The EMBL Outstation-European Bioinformatics Institute (EBI) is a center for research and services in bioinformatics. It serves researchers in molecular biology, genetics, medicine, and agriculture from academia, and the agricultural, biotechnology, chemical, and pharmaceutical industries. The Institute manages and makes available databases of biological data including nucleic acid, protein sequences, and macromolecular structures. It provides to this community bioinformatics services relevant to molecular biology free of charge over the Internet. Some of these databases and services are described in this review.

Computational Biology↗

TassDB: a database of alternative tandem splice sites.

Subtle alternative splice events at tandem splice sites are frequent in eukaryotes and substantially increase the complexity of transcriptomes and proteomes. We have developed a relational database, TassDB (TAndem Splice Site DataBase), which stores extensive data about alternative splice events at GYNGYN donors and NAGNAG acceptors. These splice events are of subtle nature since they mostly result in the insertion/deletion of a single amino acid or the substitution of one amino acid by two others. Currently, TassDB contains 114 554 tandem splice sites of eight species, 5209 of which have EST/mRNA evidence for alternative splicing. In addition, human SNPs that affect NAGNAG acceptors are annotated. The database provides a user-friendly interface to search for specific genes or for genes containing tandem splice sites with specific features as well as the possibility to download large datasets. This database should facilitate further experimental studies and large-scale bioinformatics analyses of tandem splice sites. The database is available at http://helios.informatik.uni-freiburg.de/TassDB/.

Alternative Splicing↗

Toxicogenomics in risk assessment: an overview of an HESI collaborative research program.

The value of genomic approaches in hypothesis generation is being realized as a tool for understanding toxicity and consequently contributing to an assessment of drug and chemical safety. In 1999 the membership of the International Life Sciences Institute Health and Environmental Sciences Institute formed a committee to develop a collaborative scientific program to address issues, challenges, and opportunities afforded by the emerging field of toxicogenomics. Experts and advisors from academia and government laboratories participate on the committee, along with approximately 30 corporate member organizations from the pharmaceutical, agrochemical, chemical, and consumer products industries. The committee has designed, conducted, and analyzed numerous toxicogenomic experiments within the broad fields of hepatotoxicity, nephrotoxicity, and genotoxicity. The considerable body of data generated by these programs has been instrumental in increasing understanding of sources of biological and technical variability in the alignment of toxicant-induced transcription changes with the accepted mechanism of action of these agents and the challenges in the consistent analysis and sharing of the voluminous data sets generated by these approaches. Recognizing the importance of standardized microarray data formats and public repository databases as the mechanism by which microarray data can be compared and interpreted by the scientific community, the committee has partnered with the European Bioinformatics Institute to develop a database to house the data generated by its collaborative research.

Gene Expression Profiling↗

PolyA_DB 2: mRNA polyadenylation sites in vertebrate genes.

Polyadenylation of nascent transcripts is one of the key mRNA processing events in eukaryotic cells. A large number of human and mouse genes have alternative polyadenylation sites, or poly(A) sites, leading to mRNA variants with different protein products and/or 3'-untranslated regions (3'-UTRs). PolyA_DB 2 contains poly(A) sites identified for genes in several vertebrate species, including human, mouse, rat, chicken and zebrafish, using alignments between cDNA/ESTs and genome sequences. Several new features have been added to the database since its last release, including syntenic genome regions for human poly(A) sites in seven other vertebrates and cis-element information adjacent to poly(A) sites. Trace sequences are used to provide additional evidence for poly(A/T) tails in cDNA/ESTs. The updated database is intended to broaden poly(A) site coverage in vertebrate genomes, and provide means to assess the authenticity of poly(A) sites identified by bioinformatics. The URL for this database is http://polya.umdnj.edu/PolyA_DB2.

Animals↗

MinSet: a general approach to derive maximally representative database subsets by using fragment dictionaries and its application to the SCOP database.

MOTIVATION: The size of current protein databases is a challenge for many Bioinformatics applications, both in terms of processing speed and information redundancy. It may be therefore desirable to efficiently reduce the database of interest to a maximally representative subset. RESULTS: The MinSet method employs a combination of a Suffix Tree and a Genetic Algorithm for the generation, selection and assessment of database subsets. The approach is generally applicable to any type of string-encoded data, allowing for a drastic reduction of the database size whilst retaining most of the information contained in the original set. We demonstrate the performance of the method on a database of protein domain structures encoded as strings. We used the SCOP40 domain database by translating protein structures into character strings by means of a structural alphabet and by extracting optimized subsets according to an entropy score that is based on a constant-length fragment dictionary. Therefore, optimized subsets are maximally representative for the distribution and range of local structures. Subsets containing only 10% of the SCOP structure classes show a coverage of >90% for fragments of length 1-4. AVAILABILITY: http://mathbio.nimr.mrc.ac.uk/~jkleinj/MinSet. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Algorithms↗

Bioinformatics. A user's perspective.

This review provides an overview of bioinformatics from the user's point of view. Bioinformatics, defined as the application of computers, databases, and computational methods to the management of biologic information, is essential for almost every aspect of data management in modern biology. The rapid accumulation of genomic sequence information together with the wide availability of new technologies that analyze global gene expression patterns have created an information overload. Molecular biology labs are increasingly dependent on computers, large-capacity databases, search and analysis tools, and high-quality Internet connections. Currently available bioinformatics tools are discussed and a general approach is outlined. Using the resources and approaches in this review, readers should be able to form their own view of bioinformatics and tailor the solutions to the information overload according to their needs.

Computational Biology↗

Novel developments with the PRINTS protein fingerprint database.

The PRINTS database of protein family 'fingerprints' is a diagnostic resource that complements the PROSITE dictionary of sites and patterns. Unlike regular expressions, fingerprints exploit groups of conserved motifs within sequence alignments to build characteristic signatures of family membership. Thus fingerprints inherently offer improved diagnostic reliability by virtue of the mutual context provided by motif neighbours. To date, 600 fingerprints have been constructed and stored in PRINTS, representing a 50% increase in the size of the database in the last year. The current version, 13.0, encodes approximately 3000 motifs, covering a range of globular and membrane proteins, modular polypeptides, and so on. The database is accessible via UCL's Bioinformatics World Wide Web (WWW) server at http://www.biochem.ucl.ac.uk/bsm/dbbrowser / . We describe here progress with the database, its Web interface, and a recent exciting development: the integration of a novel colour alignment editor (http://www.biochem.ucl.ac.uk/bsm/dbbrowser++ +/CINEMA ), which allows visualisation and interactive manipulation of PRINTS alignments over the Internet.

Amino Acid Sequence↗

u-Genome: a database on genome design in unicellular genomes.

Unicellular eukaryotes were among the first ones to be selected for complete genome sequencing because of the small size of their genomes and their interactions with humans and a broad range of animals and plants. Currently, ten completely sequenced unicellular genome sequences have been publicly released and as the number of available unicellular genomes increases, comparative genomics analysis within this group of organisms becomes more and more instructive. However, such an analysis is difficult to carry out without a suitable platform gathering not only the original annotations but also relevant information available in public databases or obtained by applying common bioinformatics methods. With the aim of solving these difficulties, we have developed a web-accessible database named u-Genome, the unicellular genome design database. The database is unique in featuring three datasets namely (1) orthologous proteins (2) paralogous proteins and (3) statistical distributions on exons, introns, intergenic DNA and correlations between them. A tool, Uniview, designed to visualize the gene structures for individual genes in the genome is also integrated. This database is of importance in understanding unicellular genome design and architecture and evolution related studies. The database is available through a web interface at http://sege.ntu.edu.sg/wester/ugenome.

Animals↗

Database links are a foundation for interoperability.

Several techniques are being introduced into the bioinformatics community to permit interoperation between molecular biology databases (DBs). The common factor to these approaches is the creation of links between entities in different DBs. Links can connect pieces of information about a single protein that are partitioned across multiple DBs, and can also encode relationships between different biological entities, such as relationships between an enzyme, its gene and its catalytic activity. This article provides an overview of the DB-interoperation problem, and offers several solutions. It discusses how links are used in molecular biology DBs, and describes the potential stumbling blocks when DB links are created and used.

Biotechnology↗

Progress with the PRINTS protein fingerprint database.

PRINTS is a compendium of protein motif 'fingerprints' derived from the OWL composite sequence database. Fingerprints are groups of motifs within sequence alignments whose conserved nature allows them to be used as signatures of family membership. To date, 400 fingerprints have been constructed and stored in Prints, the size of which has doubled in the last year. The current version, 9.0, encodes approximately 2000 motifs, covering a range of globular and membrane proteins, modular polypeptides, and so on. Fingerprints inherently offer improved diagnostic reliability over single motif methods by virtue of the mutual context provided by motif neighbours. PRINTS thus provides a useful adjunct to the widely used PROSITE dictionary of patterns. The database is now accessible via the Database Browser on the UCL Bioinformatics server at http://www.biochem.ucl.ac.uk/bsm/dbbrowser .

Amino Acid Sequence↗

The PRINTS database of protein fingerprints: a novel information resource for computational molecular biology.

PRINTS is a compendium of protein motif fingerprints derived from the OWL composite sequence database. Fingerprints are groups of motifs within sequence alignments whose conserved nature allows them to be used as signatures of family membership. Fingerprints inherently offer improved diagnostic reliability over single motif methods by virtue of the mutual context provided by motif neighbors. To date, 650 fingerprints have been constructed and stored in PRINTS, the size of which has doubled in the last 2 years. The current version, 14.0, encodes 3500 motifs, covering a range of globular and membrane proteins, modular polypeptides, and so on. The database is now accessible via the UCL Bioinformatics Server on http:@ www.biochem.ucl.ac.uk/bsm/dbbrowser/. We describe here progress with the database, its compilation and interrogation software, and its Web interface.

Amino Acid Sequence↗

Using proteomics and network analysis to elucidate the consequences of synaptic protein oxidation in a PS1 + AbetaPP mouse model of Alzheimer's disease.

Increasing evidence suggests that oxidative injury is involved in the pathogenesis of many age-related neurodegenerative disorders, including Alzheimer's disease (AD). Identifying the protein targets of oxidative stress is critical to determine which proteins may be responsible for the neuronal impairments and subsequent cell death that occurs in AD. In this study, we have applied a high-throughput shotgun proteomic approach to identify the targets of protein carbonylation in both aged and PS1 + AbetaPP transgenic mice. However, because of the inherent difficulties associated with proteomic database searching algorithms, several newly developed bioinformatic tools were implemented to ascertain a probability-based discernment between correct protein assignments and false identifications to improve the accuracy of protein identification. Assigning a probability to each identified peptide/protein allows one to objectively monitor the expression and relative abundance of particular proteins from diverse samples, including tissue from transgenic mice of mixed genetic backgrounds. This robust bioinformatic approach also permits the comparison of proteomic data generated by different laboratories since it is instrument- and database-independent. Applying these statistical models to our initial studies, we detected a total of 117 oxidatively modified (carbonylated) proteins, 59 of which were specifically associated with PS1 + AbetaPP mice. Pathways and network component analyses suggest that there are three major protein networks that could be potentially altered in PS1 + AbetaPP mice as a result of oxidative modifications. These pathways are 1) iNOS-integrin signaling pathway, 2) CRE/CBP transcription regulation and 3) rab-lyst vesicular trafficking. We believe the results of these studies will help establish an initial AD database of oxidatively modified proteins and provide a foundation for the design of future hypothesis driven research in the areas of aging and neurodegeneration.

Activating Transcription Factor 2↗

An interactive web-based Pseudomonas aeruginosa genome database: discovery of new genes, pathways and structures.

Using the complete genome sequence of Pseudomonas: aeruginosa PAO1, sequenced by the Pseudomonas: Genome Project (ftp://ftp.pseudomonas. com/data/pacontigs.121599), a genome database (http://pseudomonas. bit.uq.edu.au/) has been developed containing information on more than 95% of all ORFs in Pseudomonas: aeruginosa. The database is searchable by a variety of means, including gene name, position, keyword, sequence similarity and Pfam domain. Automated and manual annotation, nucleotide and peptide sequences, Pfam and SMART domains (where available), Medline and GenBank links and a scrollable, graphical representation of the surrounding genomic landscape are available for each ORF. Using the database has revealed, among other things, that P. aeruginosa contains four chemotaxis systems, two novel general secretion pathways, at least three loci encoding F17-like thin fimbriae, six novel filamentous haemagglutinin-like genes, a number of unusual composite genetic loci related to vgr/RHS: elements in Escherichia coli, a number of fix-like genes encoding a micro-oxic respiration system, novel biosynthetic pathways and 38 genes containing domains of unknown function (DUF1/DUF2). It is anticipated that this database will be a useful bioinformatic tool for the Pseudomonas: community that will continue to evolve.

Adhesins, Bacterial↗

A relational database for management of flow cytometry and ELISpot clinical trial data.

BACKGROUND: Although relational databases are widely used in bioinformatics with deposited and finalized data, they have not received widespread usage among immunologists for managing raw laboratory data such as that generated by ELISpot or flow cytometry assays. Almost no published guidance exists for immunologists to design appropriate and useful data management systems. METHODS: We describe the design and implementation of a Microsoft Access relational database used in a clinical trial in which the primary immunogenicity measures were ELISpot and intracellular cytokine staining. RESULTS: Our data management system enabled us to perform sophisticated queries and to interpret our data as quantitatively as possible. It could easily be used without modification by other researchers using automated plate reading of ELISpot plates or four color flow cytometry. CONCLUSION: We illustrate in detail the use of a flexible data management system for two of the most widely used immunological techniques. Minor modifications for more colors or other outputs can easily be implemented. Based on this example, other modifications could be easily envisaged for any other quantitative output.

Biological Specimen Banks↗

GeneInfoViz: constructing and visualizing gene relation networks.

Large amounts of knowledge about genes have been stored in public databases. One of the most challenging problems in Bioinformatics is, given all the information about the genes in the databases, determining the relationships between the genes. For example, how can we determine if genes are related and how closely they are related based on existing knowledge about their biological roles. We developed GeneInfoViz, a web tool for batch retrieval of gene information and construction and visualization of gene relation networks. We created a database containing compiled Gene Ontology information for the genes of several model organisms. Users can batch search for a group of genes and get the Gene Ontology terms that are associated with the genes. Directed acyclic graphs are generated to show the hierarchical structure of the Gene Ontology tree. GeneInfoViz calculates an adjacency matrix to determine whether the genes are related and, if so, how closely they are related based on biological processes, molecular functions, or cellular components they are associated with and then displays a dynamic graph layout of the network among the selected genes.

Databases, Genetic↗