Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Storage And Retrieval”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

Automatic processing of multilingual medical terminology: applications to thesaurus enrichment and cross-language information retrieval.

OBJECTIVES: We present in this article experiments on multi-language information extraction and access in the medical domain. For such applications, multilingual terminology plays a crucial role when working on specialized languages and specific domains. MATERIAL AND METHODS: We propose firstly a method for enriching multilingual thesauri which extracts new terms from parallel corpora, and secondly, a new approach for bilingual lexicon extraction from comparable corpora, which uses a bilingual thesaurus as a pivot. We illustrate their use in multi-language information retrieval (English/German) in the medical domains. RESULTS: Our experiments show that these automatically extracted bilingual lexicons are accurate enough (85% precision for term extraction) for semi-automatically enriching mono- or bi-lingual thesauri such as the universal medical language system, and that their use in cross-language information retrieval significantly improves the retrieval performance (from 22 to 40% average precision) and clearly outperforms existing bilingual lexicon resources (both general lexicons and specialized ones). CONCLUSION: We show in this paper first that bilingual lexicon extraction from parallel corpora in the medical domain could lead to accurate, specialized lexicons, which can be used to help enrich existing thesauri and second that bilingual lexicons extracted from comparable corpora outperform general bilingual resources for cross-language information retrieval.

Electronic Data Processing↗

GLYDE-an expressive XML standard for the representation of glycan structure.

The amount of glycomics data being generated is rapidly increasing as a result of improvements in analytical and computational methods. Correlation and analysis of this large, distributed data set requires an extensible and flexible representational standard that is also 'understood' by a wide range of software applications. An XML-based data representation standard that faithfully captures essential structural details of a glycan moiety along with additional information (such as data provenance) to aid the interpretation and usage of glycan data, will facilitate the exchange of glycomics data across the scientific community. To meet this need, we introduce GLYcan Data Exchange (GLYDE) standard as an XML-based representation format to enable interoperability and exchange of glycomics data. An online tool () for the conversion of other representations to GLYDE format has been developed.

Carbohydrate Sequence↗

VisCoSe: visualization and comparison of consensus sequences.

We introduce visualization and comparison of consensus sequences (VisCoSe) as a WWW service and a stand-alone command line Perl script for visualizing and comparing consensus sequences of protein and nucleotide sequences. VisCoSe is the only interface available that simultaneously calculates consensus sequences of multiple data sets and automatically compares these consensus sequences. Furthermore, VisCoSe allows visualization of chemical properties of amino acids.

Algorithms↗

SPIRS, WinSPIRS, and OVID: a comparison of three MEDLINE-on-CD-ROM interfaces.

Three MEDLINE-on-CD-ROM interfaces are compared: SPIRS (version 3.11) and WinSPIRS (version 1.0) from SilverPlatter and OVID (version 3.0, DOS and Windows interfaces) from CD Plus Technologies. Though the database is the same, there are substantial differences among the interfaces in the way these data are presented and can be searched. These different approaches are discussed, and a detailed comparative table is included. It is obvious that all three interfaces are quite good yet none of them is perfect; each has desirable and unfortunate features. Together, they offer an enormous range of possibilities. Users would benefit if most of the better features (e.g., easy menu, free-text retrieval, pre-exploded thesaurus terms) were implemented in future versions of these interfaces and if system operators were given greater latitude to determine the system defaults appropriate to their specific situations and customers.

CD-ROM↗

BotDB: A database resource for the clostridial neurotoxins.

BotDB is a database designed to encapsulate the rapidly expanding amount of information about the structure and function of the botulinum (BoNT) and tetanus neurotoxins and to track a variety of basic and applied research efforts. The AceDB management system was chosen for this project because of its flexibility in manipulating semistructured data sets and for its information retrieval query languages. In addition to storing amino and nucleic acid sequences of the clostridial neurotoxin genes and proteins, BotDB provides sequence data for new classes of objects, including neurotoxin mutants, substrates and their mutants, associated nontoxic proteins, and C-fragment vaccine candidates. New data types provide information on detection assays for the neurotoxins and on structural data from X-ray crystallographic and circular dichroism spectroscopic studies. Kinetic parameters from biochemical experiments include reaction rates for substrate cleavage and block of neurotransmission. The structures and kinetic characteristics of presently known chemical inhibitors are also being archived. All of these data are associated with citations of the relevant literature for on-line annotation. Graphics viewer programs are provided to display stored images and three-dimensional representations of protein structures. BotDB is in the alpha test phase of development and will become a publicly available Web site.

Amino Acid Sequence↗

Molecular Probe Data Base: a database on synthetic oligonucleotides.

The Molecular Probe Data Base (MPDB) was designed to collect and make information on synthetic oligonucleotides available on-line. This paper briefly describes its purpose, contents and structure, forms and mode of data distribution. Particular emphasis is given to recent data extension and system enhancements that have been carried out in order to simplify access to MPDB for unskilled users.

Base Sequence↗

Information management for data retrieval in a picture archive and communication system.

Data stored in a Picture archive and communication system (PACS) must be organized to permit efficient retrieval. The concept of a unique data object identifier (UID) permits a fundamental partitioning of the problem into a storage system, indexed by UID, and a database containing descriptive elements. The database serves to map user retrieval requests, expressed in terms of clinically relevant descriptive elements, and into UIDs of specific data objects. Different data organization mechanisms are employed by imaging modalities, thereby making the structure of a generic PACS database complex. One solution may be derived by analogy from film-based systems. Images and other data objects may be organized, by application of modality and site specific rules, into electronic folders. Folders may then be organized into a patient master folder or user-defined reference folders. Such an organization provides an easily understood user access model for information stored in PACS archives. This report presents a PACS architecture comprised of an Information Management System (IMS) and Information Storage System (ISS). Entity-relationship diagrams are presented to define a schema for the IMS database, based on the folder analogy. The folder concept and its relationship to the ACR-NEMA Standard are discussed.

Computer Systems↗

OligoFaktory: a visual tool for interactive oligonucleotide design.

SUMMARY: The OligoFaktory is a set of tools for the design, on an arbitrary number of target sequences, of high-quality long oligonucleotide for micro-array, of primer pair for PCR, of siRNA and more. The user-centered interface exists in two flavours: a web portal and a standalone software for Mac OS X Tiger. A unified presentation of results provides overviews with distribution charts and relative location bar graphs, as well as detailed features for each oligonucleotide. Input and output files conform to a common XML interchange file format to allow both automatic generation of input data, archiving, and post-processing of results. The design pipeline can use BLAST servers to evaluate specificity of selected oligonucleotides. AVAILABILITY: The web portal http://ueg.ulb.ac.be/oligofaktory/; the software for Macintosh: http://www.oligofaktory.org/

Algorithms↗

Associative database of protein sequences.

MOTIVATION: We present a new concept that combines data storage and data analysis in genome research, based on an associative network memory. As an illustration, 115 000 conserved regions from over 73 000 published sequences (i.e. from the entire annotated part of the SWISSPROT sequence database) were identified and clustered by a self-organizing network. Similarity and kinship, as well as degree of distance between the conserved protein segments, are visualized as neighborhood relationship on a two-dimensional topographical map. RESULTS: Such a display overcomes the restrictions of linear list processing and allows local and global sequence relationships to be studied visually. Families are memorized as prototype vectors of conserved regions. On a massive parallel machine, clustering and updating of the database take only a few seconds; a rapid analysis of incoming data such as protein sequences or ESTs is carried out on present-day workstations. AVAILABILITY: Access to the database is available at http://www.bioinf.mdc-berlin.de/unter2.html++ + CONTACT: (hanke,lehmann,reich)@mdc-berlin.de; bork@embl-heidelberg.de

Amino Acid Sequence↗

ARROGANT: an application to manipulate large gene collections.

ARROGANT (ARRay OrGANizing Tool) is a software tool developed to facilitate the identification, annotation and comparison of large collections of genes or clones. The objective is to enable users to compile gene/clone collections from different databases, allowing them to design experiments and analyze the collections as well as associated experimental data efficiently. ARROGANT can relate different sequence identifiers to their common reference sequence using the UniGene database, allowing for the comparison of data from two different microarray experiments. ARROGANT has been successfully used to analyze microarray expression data for colon cancer, to compile genes potentially related to cardiac diseases for subsequent resequencing (to identify single nucleotide polymorphisms, SNPs), to design a new comprehensive human cDNA microarray for cancer, to combine and compare expression data generated by different microarrays and to provide annotation for genes on custom and Affymetrix chips.

Base Sequence↗

WormBase as an integrated platform for the C. elegans ORFeome.

The ORFeome project has validated and corrected a large number of predicted gene models in the nematode C. elegans, and has provided an enormous resource for proteome-scale studies. To make the resource useful to the research and teaching community, it needs to be integrated with other large-scale data sets, including the C. elegans genome, cell lineage, neurological wiring diagram, transcriptome, and gene expression map. This integration is also critical because the ORFeome data sets, like other 'omics' data sets, have significant false-positive and false-negative rates, and comparison to related data is necessary to make confidence judgments in any given data point. WormBase, the central data repository for information about C. elegans and related nematodes, provides such a platform for integration. In this report, we will describe how C. elegans ORFeome data are deposited in the database, how they are used to correct gene models, how they are integrated and displayed in the context of other data sets at the WormBase Web site, and how WormBase establishes connection with the reagent-based resources at the ORFeome project Web site.

Animals↗

Visualization of genomic aberrations using Affymetrix SNP arrays.

MOTIVATION: DNA copy number aberrations are frequently found in different types of cancer. Recent developments of microarray-based approaches have broadened the knowledge on number and structure of such aberrations. High-density single nucleotide polymorphism (SNP) microarrays provide an extremely high resolution with up to 500,000 SNPs per genome. Owing to the enormous amount of data the detection of common aberrations in large datasets is a great challenge. We describe a novel open source software tool--IdeogramBrowser--which was specifically designed for use with the Affymetrix SNP arrays. It provides an interactive karyotypic visualization of multiple aberration profiles and direct links to GeneCards. Visualization of consensus regions together with gene representation allows the explorative assessment of the data. AVAILABILITY: IdeogramBrowser and its source code are freely available under a creative commons license and can be obtained from http://www.informatik.uni-ulm.de/ni/staff/HKestler/ideo/. IdeogramBrowser is a platform independent Java application.

Algorithms↗

Bias explorer: measurements of compositional bias in EMBL and GenBank sequence files.

A Windows application for compositional analysis of sequenced genomes (EMBL or GenBank flat files) is available as freeware. The application allows the user to quantify word bias using Markov chain analysis and it allows the user to generate sliding window data for GC-skew, AT-skew, purine excess, keto excess and discrete word counts. The mathematical routines reside in a dynamic link library (DLL), which can be used independently by other applications. The software is available for download at http://www.dfuni.dk/~anfu/Bioinformatics/Main.htm.

Bias↗

The MEROPS database as a protease information system.

Peptidases (often termed proteases) are of great relevance to biology, medicine, and biotechnology. This practical importance creates a need for an integrated source of information about peptidases. In the MEROPS database (www.merops.ac.uk), peptidases are classified by structural similarities in the parts of the molecules responsible for their enzymatic activity. They are grouped into families on the basis of amino acid sequence homology, and the families are assembled into clans in light of evidence that they share common ancestry. The evidence for clan-level relationships usually comes from similarities in tertiary structure, but we suggest that secondary structure profiles may also be useful in the future. The classification forms a framework around which a wealth of supplementary information about the peptidases is organized. This includes images of three-dimensional structures, alignments of matching human and mouse ESTs, comments on biomedical relevance, human and other gene symbols, and literature references linked to PubMed. For each family, there is an amino acid sequence alignment and a dendrogram. There is a list of all peptidases known from each of over 1000 species, together with summary data for the distributions of the families and clans throughout the major groups of organisms. A set of online searches provides access to information about the location of peptidases on human chromosomes and peptidase substrate specificity.

Amino Acid Sequence↗

Gene structure identification with MyGV using cDNA evidence and protein homologs to improve ab initio predictions.

UNLABELLED: MyGV is an application to visualize (potentially genome-scale) gene structure annotation and prediction. The output of any external gene prediction program can be easily converted to a generalized format for input into MyGV. The application displays all input simultaneously in graphical representation, with a toggle option for a text-based view. Zooming capabilities allow detailed comparisons for specific genome locations. The tool is particularly helpful for refinement of ab initio predicted gene structures by spliced alignment with cDNA or protein homologs. AVAILABILITY: The program was written in Java and is freely available to non-commercial users by electronic download from http://bioinformatics.iastate.edu/bioinformatics2go/MyGV.

Animals↗

OpScan-MIMS. A data network system.

The expansion of a unique system for medical, epidemiological, and laboratory information management is described. The system utilizes an automated, optically scanned data entry mode (OpScan), and a generalized, interactive storage and retrieval software package (MIMS). Aggregate medical, epidemiological, and laboratory information are monitored, with the permission of the locality, via telephone terminals at four sexually transmitted disease clinics. OpScan-MIMS combines an efficient, rapid, and cost-effective method for data entry with a user-friendly computer system that requires no specialized training in computer language.

Centers for Disease Control and Prevention, U.S.↗