Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Storage And Retrieval”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

EMBL-Align: a new public nucleotide and amino acid multiple sequence alignment database.

UNLABELLED: The submission of multiple sequence alignment data to EMBL has grown 30-fold in the past 10 years, creating a problem of archiving them. The EBI has developed a new public database of multiple sequence alignments called EMBL-Align. It has a dedicated web-based submission tool, Webin-Align. Together they represent a comprehensive data management solution for alignment data. Webin-Align accepts all the common alignment formats and can display data in CLUSTALW format as well as a new standard EMBL-Align flat file format. The alignments are stored in the EMBL-Align database and can be queried from the EBI SRS (Sequence Retrieval System) server. AVAILABILITY: Webin-Align: http://www.ebi.ac.uk/embl/Submission/align_top.html, EMBL-Align: ftp://ftp.ebi.ac.uk/pub/databases/embl/align, http://srs.ebi.ac.uk/

Amino Acid Sequence↗

ProLysED: an integrated database and meta-server of bacterial protease systems.

UNLABELLED: Bacterial proteases are an important group of enzymes that have very diverse biochemical and cellular functions. Proteases from prokaryotic sources also have a wide range of uses, either in medicine as pathogenic factors or in industry and therapeutics. ProLysED (Prokaryotic Lysis Enzymes Database), our meta-server integrated database of bacterial proteases, is a useful, albeit very niche, resource. The features include protease classification browsing and searching, organism-specific protease browsing, molecular information and visualisation of protease structures from the Protein Data Bank (PDB) as well as predicted protease structures. AVAILABILITY: ProLysED is integrated into the ProLysES (Prokaryotic Lysis Enzymes Site) website at http://genome.ukm.my/prolyses/. Access to the ProLysED database is free for academic users upon registration.

Amino Acid Sequence↗

The Enhanced Microbial Genomes Library.

Since the obtention of the complete sequence of Haemophilus influenzae Rd in 1995, the number of bacterial genomes entirely sequenced has regularly increased. A problem is that the quality of the annotations of these very large sequences is usually lower than those of the shorter entries encountered in the repository collections. Moreover, classical sequence database management systems have difficulties in handling entries of that size. In this context, we have decided to build the Enhanced Microbial Genomes Library (EMGLib) in which these two problems are alleviated. This library contains all the complete genomes from bacteria already sequenced and the yeast genome in GenBank format. The annotations are improved by the introduction of data on codon usage, gene orientation on the chromosome and gene families. It is possible to access EMGLib through two database systems set up on World Wide Web servers: the PBIL server at http://pbil.univ-lyon1.fr/emglib/emglib. html and the MICADO server at http://locus.jouy.inra.fr/micado

Base Sequence↗

The distribution of RNA motifs in natural sequences.

Functional analysis of genome sequences has largely ignored RNA genes and their structures. We introduce here the notion of 'ribonomics' to describe the search for the distribution of and eventually the determination of the physiological roles of these RNA structures found in the sequence databases. The utility of this approach is illustrated here by the identification in the GenBank database of RNA motifs having known binding or chemical activity. The frequency of these motifs indicates that most have originated from evolutionary drift and are selectively neutral. On the other hand, their distribution among species and their location within genes suggest that the destiny of these motifs may be more elaborate. For example, the hammerhead motif has a skewed organismal presence, is phylogenetically stable and recent work on a schistosome version confirms its in vivo biological activity. The under-representation of the valine-binding motif and the Rev-binding element in GenBank hints at a detrimental effect on cell growth or viability. Data on the presence and the location of these motifs may provide critical guidance in the design of experiments directed towards the understanding and the manipulation of RNA complexes and activities in vivo.

Adenosine Triphosphate↗

PDBML: the representation of archival macromolecular structure data in XML.

SUMMARY: The Protein Data Bank (PDB) has recently released versions of the PDB Exchange dictionary and the PDB archival data files in XML format collectively named PDBML. The automated generation of these XML files is driven by the data dictionary infrastructure in use at the PDB. The correspondences between the PDB dictionary and the XML schema metadata are described as well as the XML representations of PDB dictionaries and data files.

Amino Acid Sequence↗

DICOM modality worklist: an essential component in a PACS environment.

The development and acceptance of the digital communication in medicine (DICOM) standard has become a basic requirement for the implementation of electronic imaging in radiology. DICOM is now evolving to provide a standard for electronic communication between radiology and other parts of the hospital enterprise. In a completely integrated filmless radiology department, there are 3 core computer systems, the picture archiving and communication system (PACS), the hospital or radiology information system (HIS, RIS), and the acquisition modality. Ideally, each would have bidirectional communication with the other 2 systems. At a minimum, a PACS must be able to receive and acknowledge receipt of image and demographic data from the modalities. Similarly, the modalities must be able to send images and demographic data to the PACS. Now that basic DICOM communication protocols for query or retrieval, storage, and print classes have become established through both conformance statements and intervendor testing, there has been an increase in interest in enhancing the functionality of communication between the 3 computers. Historically, demographic data passed to the PACS have been generated manually at the modality despite the existence of the same data on the HIS or RIS. In more current sophisticated implementations, acquisition modalities are able to receive patient and study-related data from the HIS or RIS. DICOM Modality Worklist is the missing electronic link that transfers this critical information between the acquisition modalities and the HIS or RIS. This report describes the concepts, issues, and impact of DICOM Modality Worklist implementation in a PACS environment.

Forms and Records Control↗

The binding interface database (BID): a compilation of amino acid hot spots in protein interfaces.

SUMMARY: To make information about protein interactive function easily accessible, we are mining the primary scientific literature for detailed data about protein interfaces. The Binding Interface Database (BID) organizes the vast amount of protein interaction information into tables, graphical contact maps and descriptive functional profiles. Currently data on 170 interacting protein pairs are available with over 1300 mutations described. AVAILABILITY: The BID database is freely available at http://tsailab.org/BID/ To have your protein of interest entered, contact Tiffany Fischer (tiffbrink@neo.tamu.edu) or Jerry Tsai at the email below

Amino Acid Sequence↗

Laboratory Information Management Software for genotyping workflows: applications in high throughput crop genotyping.

BACKGROUND: With the advances in DNA sequencer-based technologies, it has become possible to automate several steps of the genotyping process leading to increased throughput. To efficiently handle the large amounts of genotypic data generated and help with quality control, there is a strong need for a software system that can help with the tracking of samples and capture and management of data at different steps of the process. Such systems, while serving to manage the workflow precisely, also encourage good laboratory practice by standardizing protocols, recording and annotating data from every step of the workflow. RESULTS: A laboratory information management system (LIMS) has been designed and implemented at the International Crops Research Institute for the Semi-Arid Tropics (ICRISAT) that meets the requirements of a moderately high throughput molecular genotyping facility. The application is designed as modules and is simple to learn and use. The application leads the user through each step of the process from starting an experiment to the storing of output data from the genotype detection step with auto-binning of alleles; thus ensuring that every DNA sample is handled in an identical manner and all the necessary data are captured. The application keeps track of DNA samples and generated data. Data entry into the system is through the use of forms for file uploads. The LIMS provides functions to trace back to the electrophoresis gel files or sample source for any genotypic data and for repeating experiments. The LIMS is being presently used for the capture of high throughput SSR (simple-sequence repeat) genotyping data from the legume (chickpea, groundnut and pigeonpea) and cereal (sorghum and millets) crops of importance in the semi-arid tropics. CONCLUSION: A laboratory information management system is available that has been found useful in the management of microsatellite genotype data in a moderately high throughput genotyping laboratory. The application with source code is freely available for academic users and can be downloaded from http://www.icrisat.org/gt-bt/lims/lims.asp.

Algorithms↗

Paper2sequences: retrieval of sequences listed in a publication.

Our web-based tool simplifies the often laborious procedure of retrieving a set of biosequences in a publication or webpage. As a front-end to the Bioperl toolkit, it accepts as an input a list of identifiers. They are specified in an ASCII table (copy-pasted from the publication's PDF or HTML page) and give rise to queries in multiple databases for the protein/nucleic acid data specified. Currently, GenBank, PIR (Protein Information Resource) and Swiss-Prot are supported. For any sequence accession code listed, the database can be specified and, if retrieval fails, automatic lookup for the same code in other databases can be requested. Sequence length information (if specified) and heuristic rules are used to drive the lookup if multiple protein coding sequences (CDS) are part of a single accession. Warnings are issued in cases of ambiguities and inconsistencies. An advanced option enables the user to format the output in whatever format they wish.

Amino Acid Sequence↗

PEDE (Pig EST Data Explorer): construction of a database for ESTs derived from porcine full-length cDNA libraries.

We generated the PEDE (Pig EST Data Explorer; http://pede.dna.affrc.go.jp/) database using sequences assembled from porcine 5' ESTs from oligo-capped full-length cDNA libraries. Thus far we have performed EST analysis of various organs (thymus, spleen, uterus, lung, liver, ovary and peripheral blood mononuclear cells) and assembled 68,076 high-quality sequences into 5546 contigs and 28,461 singlets. PEDE provides a search interface for getting results of homology searches and enables users to obtain information on sequence data and cDNA clones of interest. Single-nucleotide polymorphisms detected through comparison of the EST sequences are classified by origin (western and oriental breeds) and are searchable in the database. This database system can accelerate analyses of livestock traits and yields information that can lead to new applications in pigs as model systems for medical research.

Animals↗

SNP Function Portal: a web database for exploring the function implication of SNP alleles.

MOTIVATION: Finding the potential functional significance of SNPs is a major bottleneck in understanding genome-wide SNP scanning results, as the related functional data are distributed across many different databases. The SNP Function Portal is designed to be a clearing house for all public domain SNP functional annotation data, as well as in-house functional annotations derived from different data sources. It currently contains SNP functional annotations in six major categories including genomic elements, transcription regulation, protein function, pathway, disease and population genetics. Besides extensive SNP functional annotations, the SNP Function Portal includes a powerful search engine that accepts different types of genetic markers as input and identifies all genetically related SNPs based on the HapMap Phase II data as well as the relationship of different markers to known genes. As a result, our system allows users to identify the potential biological impact of genetic markers and complex relationships among genetic markers and genes, and it greatly facilitates knowledge discovery in genome-wide SNP scanning experiments. AVAILABILITY: http://brainarray.mbni.med.umich.edu/Brainarray/Database/SearchSNP/snpfunc.aspx.

Alleles↗

Functional annotation by identification of local surface similarities: a novel tool for structural genomics.

BACKGROUND: Protein function is often dependent on subsets of solvent-exposed residues that may exist in a similar three-dimensional configuration in non homologous proteins thus having different order and/or spacing in the sequence. Hence, functional annotation by means of sequence or fold similarity is not adequate for such cases. RESULTS: We describe a method for the function-related annotation of protein structures by means of the detection of local structural similarity with a library of annotated functional sites. An automatic procedure was used to annotate the function of local surface regions. Next, we employed a sequence-independent algorithm to compare exhaustively these functional patches with a larger collection of protein surface cavities. After tuning and validating the algorithm on a dataset of well annotated structures, we applied it to a list of protein structures that are classified as being of unknown function in the Protein Data Bank. By this strategy, we were able to provide functional clues to proteins that do not show any significant sequence or global structural similarity with proteins in the current databases. CONCLUSION: This method is able to spot structural similarities associated to function-related similarities, independently on sequence or fold resemblance, therefore is a valuable tool for the functional analysis of uncharacterized proteins. Results are available at http://cbm.bio.uniroma2.it/surface/structuralGenomics.html.

Algorithms↗

GeneTools--application for functional annotation and statistical hypothesis testing.

BACKGROUND: Modern biology has shifted from "one gene" approaches to methods for genomic-scale analysis like microarray technology, which allow simultaneous measurement of thousands of genes. This has created a need for tools facilitating interpretation of biological data in "batch" mode. However, such tools often leave the investigator with large volumes of apparently unorganized information. To meet this interpretation challenge, gene-set, or cluster testing has become a popular analytical tool. Many gene-set testing methods and software packages are now available, most of which use a variety of statistical tests to assess the genes in a set for biological information. However, the field is still evolving, and there is a great need for "integrated" solutions. RESULTS: GeneTools is a web-service providing access to a database that brings together information from a broad range of resources. The annotation data are updated weekly, guaranteeing that users get data most recently available. Data submitted by the user are stored in the database, where it can easily be updated, shared between users and exported in various formats. GeneTools provides three different tools: i) NMC Annotation Tool, which offers annotations from several databases like UniGene, Entrez Gene, SwissProt and GeneOntology, in both single- and batch search mode. ii) GO Annotator Tool, where users can add new gene ontology (GO) annotations to genes of interest. These user defined GO annotations can be used in further analysis or exported for public distribution. iii) eGOn, a tool for visualization and statistical hypothesis testing of GO category representation. As the first GO tool, eGOn supports hypothesis testing for three different situations (master-target situation, mutually exclusive target-target situation and intersecting target-target situation). An important additional function is an evidence-code filter that allows users, to select the GO annotations for the analysis. CONCLUSION: GeneTools is the first "all in one" annotation tool, providing users with a rapid extraction of highly relevant gene annotation data for e.g. thousands of genes or clones at once. It allows a user to define and archive new GO annotations and it supports hypothesis testing related to GO category representations. GeneTools is freely available through www.genetools.no

Algorithms↗

Hairpins in a Haystack: recognizing microRNA precursors in comparative genomics data.

UNLABELLED: Recently, genome-wide surveys for non-coding RNAs have provided evidence for tens of thousands of previously undescribed evolutionary conserved RNAs with distinctive secondary structures. The annotation of these putative ncRNAs, however, remains a difficult problem. Here we describe an SVM-based approach that, in conjunction with a non-stringent filter for consensus secondary structures, is capable of efficiently recognizing microRNA precursors in multiple sequence alignments. The software was applied to recent genome-wide RNAz surveys of mammals, urochordates, and nematodes. AVAILABILITY: The program RNAmicro is available as source code and can be downloaded from http://www.bioinf.uni-leipzig/Software/RNAmicro.

Algorithms↗

PseudoBase: structural information on RNA pseudoknots.

PseudoBase is a database containing structural, functional and sequence data related to RNA pseudo-knots. It can be reached at http://wwwbio.LeidenUniv.nl/ approximately Batenburg/PKB.html. For each pseudoknot, thirteen items are stored, for example the relevant sequence, the stem positions of the pseudoknot, the EMBL accession number of the sequence and the support that can be given regarding the reliability of the pseudo-knot. Since the last publication, information on sizes of the stems and the loops in the pseudoknots has been added. Also added are alternative entries that produce surveys of where the pseudoknots are, sorted according to stem size or loop size.

Base Sequence↗

A new way for multidimensional medical data management: volume of interest (VOI)-based retrieval of medical images with visual and functional features.

The advances in digital medical imaging and storage in integrated databases are resulting in growing demands for efficient image retrieval and management. Content-based image retrieval (CBIR) refers to the retrieval of images from a database, using the visual features derived from the information in the image, and has become an attractive approach to managing large medical image archives. In conventional CBIR systems for medical images, images are often segmented into regions which are used to derive two-dimensional visual features for region-based queries. Although such approach has the advantage of including only relevant regions in the formulation of a query, medical images that are inherently multidimensional can potentially benefit from the multidimensional feature extraction which could open up new opportunities in visual feature extraction and retrieval. In this study, we present a volume of interest (VOI) based content-based retrieval of four-dimensional (three spatial and one temporal) dynamic PET images. By segmenting the images into VOIs consisting of functionally similar voxels (e.g., a tumor structure), multidimensional visual and functional features were extracted and used as region-based query features. A prototype VOI-based functional image retrieval system (VOI-FIRS) has been designed to demonstrate the proposed multidimensional feature extraction and retrieval. Experimental results show that the proposed system allows for the retrieval of related images that constitute similar visual and functional VOI features, and can find potential applications in medical data management, such as to aid in education, diagnosis, and statistical analysis.

Algorithms↗