SEARCH · Search PubMed
Results for “bioinformatic database”
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Changing the rules? The agreement between Celera and Science magazine concerning Celera's publication of its human genome sequence is upsetting many researchers in bioinformatics.
Explore the source record for details and available documents.
Database of p53 gene somatic mutations in human tumors and cell lines: updated compilation and future prospects.
In recent years, there has been an exponential increase in the number of p53 mutations identified in human cancers. The p53 mutation database consists of a list of point mutations in thep53 gene of human tumors and cell lines, compiled from the published literature and made available through electronic media. The database is now maintained at the International Agency for Research on Cancer (IARC) and is updated twice a year. The current version contains records on 5091 published mutations and is expected to surpass the 6000 mark in the January 1997 release. The database is available in various formats through the European Bioinformatics Institute (EBI) ftp server at: ftp://ftp.ebi.ac.uk/pub/databases/p53/ or by request from IARC (p53database@iarc.fr) and will be searchable through the SRS system in the near future. This report provides a description of the criteria for inclusion of data and of the current formats, a summary of the relevance ofp53 mutation analysis to clinical and biological questions, and a brief discussion of the prospects for future developments.
Basic immunology, allergen prediction, and bioinformatics.
Explore the source record for details and available documents.
A two-way bioinformatic street.
Explore the source record for details and available documents.
A project for ocular bioinformatics: NEIBank.
Explore the source record for details and available documents.
[Bioinformatics].
Explore the source record for details and available documents.
Genethics. Toward striking a balance in bioinformatics.
Explore the source record for details and available documents.
A web-based tool to retrieve human genome polymorphisms from public databases.
Single Nucleotide Polymorphisms (SNPs) are the most important source of variation in our genome, and an invaluable tool in the hands of researchers who investigate genetic diseases. Databases of SNPs are growing at a very fast rate, and the ability to perform large-scale, high-resolution association studies is quickly becoming a reality. In this paper we describe SNPper, a web-based tool to search for SNPs in public databases. The system allows searching for all SNPs in a given set of genes (for candidate gene studies) or in a specified region of a chromosome. The information displayed for each gene or each SNP is fully annotated and linked to the leading bioinformatics web sites. The first release of SNPper is available on the web, and has received positive feedback from the genetic and bioinformatics community.
404 not found: the stability and persistence of URLs published in MEDLINE.
MOTIVATION: The advent of the World Wide Web has enabled unprecedented supplementation of traditional journal publications, allowing access to resources, such as video, sound, software, databases, datasets too large to publish, and even supplementary information and discussion. However, unlike traditional publications, continued availability of these online resources is not guaranteed. An automated survey was conducted to quantify the growth in Uniform Resource Locators (URLs) published to date in MEDLINE abstracts, their current availability and distribution by journal. RESULTS: Of 1630 unique URLs identified, formatting and/or spelling errors were detected within 201 (12%) of them as published. After corrections were made, a survey revealed that approximately 63% of these URLs were consistently available, and another 19% were available intermittently. The rate of failure was far worse for anonymous login to FTP sites, with only 12 of 33 sites (36%) responding. This survey also shows that journals vary disproportionately in the number of web citations published, suggesting policy implementation among a few could have a profound impact overall. Out of the 306 journals with a URL published in an abstract, Bioinformatics published the most (12% of total). AVAILABILITY: URL database and program available by request.
GNARE: automated system for high-throughput genome analysis with grid computational backend.
Recent progress in genomics and experimental biology has brought exponential growth of the biological information available for computational analysis in public genomics databases. However, applying the potentially enormous scientific value of this information to the understanding of biological systems requires computing and data storage technology of an unprecedented scale. The Grid, with its aggregated and distributed computational and storage infrastructure, offers an ideal platform for high-throughput bioinformatics analysis. To leverage this we have developed the Genome Analysis Research Environment (GNARE)--a scalable computational system for the high-throughput analysis of genomes, which provides an integrated database and computational backend for data-driven bioinformatics applications. GNARE efficiently automates the major steps of genome analysis including acquisition of data from multiple genomic databases; data analysis by a diverse set of bioinformatics tools; and storage of results and annotations. High-throughput computations in GNARE are performed using distributed heterogeneous Grid computing resources such as Grid2003, TeraGrid, and the DOE Science Grid. Multi-step genome analysis workflows involving massive data processing, the use of application-specific tools and algorithms and updating of an integrated database to provide interactive web access to results are all expressed and controlled by a "virtual data" model which transparently maps computational workflows to distributed Grid resources. This paper describes how Grid technologies such as Globus, Condor, and the Gryphyn Virtual Data System were applied in the development of GNARE. It focuses on our approach to Grid resource allocation and to the use of GNARE as a computational framework for the development of bioinformatics applications.
LSAT: learning about alternative transcripts in MEDLINE.
MOTIVATION: Generation of alternative transcripts from the same gene is an important biological event due to their contribution in creating functional diversity in eukaryotes. In this work, we choose the task of extracting information around this complex topic using a two-step procedure involving machine learning and information extraction. RESULTS: In the first step, we trained a classifier that inductively learns to identify sentences about physiological transcript diversity from the MEDLINE abstracts. Using a large hand-built corpus, we compared the sentence classification performance of various text categorization methods. Support vector machines (SVMs) followed by the maximum entropy classifier outperformed other methods for the sentence classification task. The SVM with the radial basis function kernel and optimized parameters achieved Fbeta-measure of 91% during the 4-fold cross validation and of 74% when applied to all sentences in more than 12 million abstracts of MEDLINE. In the second step, we identified eight frequently present semantic categories in the sentences and performed a limited amount of semantic role labeling. The role labeling step also achieved very high Fbeta-measure for all eight categories. AVAILABILITY: The results of our two-step procedure are summarized in the LSAT database of alternative transcripts. LSAT is available at http://www.bork.embl.de/LSAT CONTACT: shah@embl.de SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Identification of key genes related to bone metastasis of breast cancer using bioinformatics methods and construction of a prognostic model.
Breast cancer (BC) ranks among the most prevalent cancers in females, with bone metastasis significantly compromising patients' quality of life and survival rates. Enhancing our comprehension of BC bone metastasis mechanisms at the molecular level holds promise for improving BC treatment and prognosis. Leveraging bioinformatics tools, we integrated multiple datasets, conducted comprehensive analyses across various databases, identified biomarkers associated with BC bone metastasis, and constructed a prognostic model. Firstly, 3 BC bone metastasis-related datasets were downloaded from gene expression omnibus, the data were merged, and batch effects were removed, followed by identification of differentially expressed genes (DEGs). Gene ontology and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analyses were performed on the DEGs. A protein-protein interaction network was constructed using the STRING database to screen hub genes. Then, survival analysis of hub genes was performed using the Cancer Genome Atlas (TCGA) database. A prognostic model was constructed using key genes with survival differences, and the model was evaluated. Two hundred ninety-two DEGs were identified. Gene ontology and KEGG pathway enrichment analysis yielded 769 biological processes (BPs), 78 cellular components, 43 molecular functions, and 50 KEGG pathways. Fifteen hub genes were selected from the protein-protein interaction network. Survival analysis revealed 6 genes related to BC survival. The prognostic model identified 4 genes with important predictive value for BC prognosis. Our study utilized bioinformatics analysis to identify a series of DEGs related to BC bone metastasis. Based on further selection of hub genes, we constructed a relatively ideal prognostic model for BC, and identified 4 genes (DLGAP5, TPX2, PLK1, and CENPN) with valuable predictive value for BC prognosis.
A database of unique protein sequence identifiers for proteome studies.
In proteome studies, identification of proteins requires searching protein sequence databases. The public protein sequence databases (e.g., NCBInr, UniProt) each contain millions of entries, and private databases add thousands more. Although much of the sequence information in these databases is redundant, each database uses distinct identifiers for the identical protein sequence and often contains unique annotation information. Users of one database obtain a database-specific sequence identifier that is often difficult to reconcile with the identifiers from a different database. When multiple databases are used for searches or the databases being searched are updated frequently, interpreting the protein identifications and associated annotations can be problematic. We have developed a database of unique protein sequence identifiers called Sequence Globally Unique Identifiers (SEGUID) derived from primary protein sequences. These identifiers serve as a common link between multiple sequence databases and are resilient to annotation changes in either public or private databases throughout the lifetime of a given protein sequence. The SEGUID Database can be downloaded (http://bioinformatics.anl.gov/SEGUID/) or easily generated at any site with access to primary protein sequence databases. Since SEGUIDs are stable, predictions based on the primary sequence information (e.g., pI, Mr) can be calculated just once; we have generated approximately 500 different calculations for more than 2.5 million sequences. SEGUIDs are used to integrate MS and 2-DE data with bioinformatics information and provide the opportunity to search multiple protein sequence databases, thereby providing a higher probability of finding the most valid protein identifications.
NEOBASE: databasing the neocortical microcircuit.
Mammals adapt to a rapidly changing world because of the sophisticated perceptual and cognitive function enabled by the neocortex. The neocortex, which has expanded to constitute nearly 80% of the human brain seems to have arisen from repeated duplication of a stereotypical template of neurons and synaptic circuits with subtle specializations in different brain regions and species. Determining the design and function of this microcircuitry is therefore of paramount importance to understanding normal and abnormal higher brain function. Recent advances in recording synaptically-coupled neurons has allowed rapid dissection of the neocortical microcircuitry thus yielding a massive amount of quantitative anatomical, electrical and gene expression data on the neurons and the synaptic circuits that connect the neurons. Due to the availability of the above mentioned data, it has now become imperative to database the neurons of the microcircuit and their synaptic connections. The NEOBASE project, aims to archive the neocortical microcircuit data in a manner that facilitates development of advanced data mining applications, statistical and bioinformatics analyses tools, custom microcircuit builders, and visualization and simulation applications. The database architecture is based on ROOT, a software environment that allows the construction of an object oriented database with numerous relational capabilities. The proposed architecture allows construction of a database that closely mimics the architecture of the real microcircuit, which facilitates the interface with virtually any application, allows for data format evolution, and aims for full interoperability with other databases. NEOBASE will provide an important resource and research tool for studying the microcircuit basis of normal and abnormal neocortical function. The database will be available to local as well as remote users using Grid based tools and technologies.
The World-Wide Web: an interface between research and teaching in bioinformatics.
The rapid expansion occurring in World-Wide Web activity is beginning to make the concepts of 'global hypermedia' and 'universal document readership realistic objectives of the new revolution in information technology. One consequence of this increase in usage is that educators and students are becoming more aware of the diversity of the knowledge base which can be accessed via the Internet. Although computerised databases and information services have long played a key role in bioinformatics these same resources can also be used to provide core materials for teaching and learning. The large datasets and archives that have been compiled for biomedical research can be enhanced with the addition of a variety of multimedia elements (images, digital videos, animation etc.). The use of this digitally stored information in structured and self-directed learning environments is likely to increase as activity across World-Wide Web increases.
YAdumper: extracting and translating large information volumes from relational databases to structured flat files.
Downloading the information stored in relational databases into XML and other flat formats is a common task in bioinformatics. This periodical dumping of information requires considerable CPU time, disk and memory resources. YAdumper has been developed as a purpose-specific tool to deal with the integral structured information download of relational databases. YAdumper is a Java application that organizes database extraction following an XML template based on an external Document Type Declaration. Compared with other non-native alternatives, YAdumper substantially reduces memory requirements and considerably improves writing performance.
The bioinformatics resource for oral pathogens.
Complete genomic sequences of several oral pathogens have been deciphered and multiple sources of independently annotated data are available for the same genomes. Different gene identification schemes and functional annotation methods used in these databases present a challenge for cross-referencing and the efficient use of the data. The Bioinformatics Resource for Oral Pathogens (BROP) aims to integrate bioinformatics data from multiple sources for easy comparison, analysis and data-mining through specially designed software interfaces. Currently, databases and tools provided by BROP include: (i) a graphical genome viewer (Genome Viewer) that allows side-by-side visual comparison of independently annotated datasets for the same genome; (ii) a pipeline of automatic data-mining algorithms to keep the genome annotation always up-to-date; (iii) comparative genomic tools such as Genome-wide ORF Alignment (GOAL); and (iv) the Oral Pathogen Microarray Database. BROP can also handle unfinished genomic sequences and provides secure yet flexible control over data access. The concept of providing an integrated source of genomic data, as well as the data-mining model used in BROP can be applied to other organisms. BROP can be publicly accessed at http://www.brop.org.