Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Storage And Retrieval”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Saccharomyces Genome Database (SGD) provides tools to identify and analyze sequences from Saccharomyces cerevisiae and related sequences from other organisms.

The Saccharomyces Genome Database (SGD; http://www.yeastgenome.org/), a scientific database of the molecular biology and genetics of the yeast Saccharomyces cerevisiae, has recently developed several new resources that allow the comparison and integration of information on a genome-wide scale, enabling the user not only to find detailed information about individual genes, but also to make connections across groups of genes with common features and across different species. The Fungal Alignment Viewer displays alignments of sequences from multiple fungal genomes, while the Sequence Similarity Query tool displays PSI-BLAST alignments of each S.cerevisiae protein with similar proteins from any species whose sequences are contained in the non-redundant (nr) protein data set at NCBI. The Yeast Biochemical Pathways tool integrates groups of genes by their common roles in metabolism and displays the metabolic pathways in a graphical form. Finally, the Find Chromosomal Features search interface provides a versatile tool for querying multiple types of information in SGD.

Amino Acid Sequence↗

Searching sequence databases via de novo peptide sequencing by tandem mass spectrometry.

There are many computer programs that can match tandem mass spectra of peptides to database-derived sequences; however, situations can arise where mass spectral data cannot be correlated with any database sequence. In such cases, sequences can be automatically deduced de novo, without recourse to sequence databases, and the resulting peptide sequences can be used to perform homologous nonexact searches of sequence databases. This article describes details on how to implement both a de novo sequencing program called "Lutefisk," and a version of FASTA that has been modified to account for sequence ambiguities inherent in tandem mass spectrometry data.

Algorithms↗

RADARS, a bioinformatics solution that automates proteome mass spectral analysis, optimises protein identification, and archives data in a relational database.

RADARS, a rapid, automated, data archiving and retrieval software system for high-throughput proteomic mass spectral data processing and storage, is described. The majority of mass spectrometer data files are compatible with RADARS, for consistent processing. The system automatically takes unprocessed data files, identifies proteins via in silico database searching, then stores the processed data and search results in a relational database suitable for customized reporting. The system is robust, used in 24/7 operation, accessible to multiple users of an intranet through a web browser, may be monitored by Virtual Private Network, and is secure. RADARS is scalable for use on one or many computers, and is suited to multiple processor systems. It can incorporate any local database in FASTA format, and can search protein and DNA databases online. A key feature is a suite of visualisation tools (many available gratis), allowing facile manipulation of spectra, by hand annotation, reanalysis, and access to all procedures. We also described the use of Sonar MS/MS, a novel, rapid search engine requiring 40 MB RAM per process for searches against a genomic or EST database translated in all six reading frames. RADARS reduces the cost of analysis by its efficient algorithms: Sonar MS/MS can identifiy proteins without accurate knowledge of the parent ion mass and without protein tags. Statistical scoring methods provide close-to-expert accuracy and brings robust data analysis to the non-expert user.

Amino Acid Sequence↗

DDBASE2.0: updated domain database with improved identification of structural domains.

MOTIVATION: Although many methods are available for the identification of structural domains from protein three-dimensional structures, accurate definition of protein domains and the curation of such data for a large number of proteins are often possible only after manual intervention. The availability of domain definitions for protein structural entries is useful for the sequence analysis of aligned domains, structure comparison, fold recognition procedures and understanding protein folding, domain stability and flexibility. RESULTS: We have improved our method of domain identification starting from the concept of clustering secondary structural elements, but with an intention of reducing the number of discontinuous segments in identified domains. The results of our modified and automatic approach have been compared with the domain definitions from other databases. On a test data set of 55 proteins, this method acquires high agreement (88%) in the number of domains with the crystallographers' definition and resources such as SCOP, CATH, DALI, 3Dee and PDP databases. This method also obtains 98% overlap score with the other resources in the definition of domain boundaries of the 55 proteins. We have examined the domain arrangements of 4592 non-redundant protein chains using the improved method to include 5409 domains leading to an update of the structural domain database. AVAILABILITY: The latest version of the domain database and online domain identification methods are available from http://www.ncbs.res.in/~faculty/mini/ddbase/ddbase.html SUPPLEMENTARY INFORMATION: http://www.ncbs.res.in/~faculty/mini/ddbase/supplementary/supplementary.html

Amino Acid Sequence↗

Automatic extraction of gene and protein synonyms from MEDLINE and journal articles.

Genes and proteins are often associated with multiple names, and more names are added as new functional or structural information is discovered. Because authors often alternate between these synonyms, information retrieval and extraction benefits from identifying these synonymous names. We have developed a method to extract automatically synonymous gene and protein names from MEDLINE and journal articles. We first identified patterns authors use to list synonymous gene and protein names. We developed SGPE (for synonym extraction of gene and protein names), a software program that recognizes the patterns and extracts from MEDLINE abstracts and full-text journal articles candidate synonymous terms. SGPE then applies a sequence of filters that automatically screen out those terms that are not gene and protein names. We evaluated our method to have an overall precision of 71% on both MEDLINE and journal articles, and 90% precision on the more suitable full-text articles alone

Electronic Data Processing↗

QGB: a system for querying sequence database fields and features.

We have developed a general system, QGB, for performing complex queries on the information in the DDBJ/EMBL/GenBank databases, including queries over the structural features of sequences implied in the FEATURE TABLE. Queries are formed in a Structured Query Language (SQL)-like syntax with language extensions to support complex types (e.g., sets, ordered sets, and records) appropriate for representing and querying sequence data. A novel aspect of QGB is its ability to deduce missing features and infer relationships among features as a consequence of constructing a parse tree of sequence structure from information described in the FEATURE TABLE. The grammar for the parse tree is implemented in a customized form of the Definite Clause Grammar syntax of the logic programming language Prolog. The logic grammar formalism was chosen because it provides a perspicuous representation for features and constraints, and Prolog provides an execution model for the grammar rules. Construction of the parse tree also identifies inconsistencies and errors in the FEATURE TABLE that can in some cases be corrected automatically and used to generate an augmented version of the table.

Base Sequence↗

Base-rate effects in category learning: a comparison of parallel network and memory storage-retrieval models.

Exemplar-memory and adaptive network models were compared in application to category learning data, with special attention to base rate effects on learning and transfer performance. Subjects classified symptom charts of hypothetical patients into disease categories, with informative feedback on learning trials and with the feedback either given or withheld on test trials that followed each fourth of the learning series. The network model proved notably accurate and uniformly superior to the exemplar model in accounting for the detailed course of learning; both the parallel, interactive aspect of the network model and its particular learning algorithm contribute to this superiority. During learning, subjects' performance reflected both category base rates and feature (symptom) probabilities in a nearly optimal manner, a result predicted by both models, though more accurately by the network model. However, under some test conditions, the data showed substantial base-rate neglect, in agreement with Gluck and Bower (1988b).

Adult↗

3D-GENOMICS: a database to compare structural and functional annotations of proteins between sequenced genomes.

The 3D-GENOMICS database (http://www.sbg.bio. ic.ac.uk/3dgenomics/) provides structural annotations for proteins from sequenced genomes. In August 2003 the database included data for 93 proteomes. The annotations stored in the database include homologous sequences from various sequence databases, domains from SCOP and Pfam, patterns from Prosite and other predicted sequence features such as transmembrane regions and coiled coils. In addition to annotations at the sequence level, several precomputed cross- proteome comparative analyses are available based on SCOP domain superfamily composition. Annotations are available to the user via a web interface to the database. Multiple points of entry are available so that a user is able to: (i) directly access annotations for a single protein sequence via keywords or accession codes, (ii) examine a sequence of interest chosen from a summary of annotations for a particular proteome, or (iii) access precomputed frequency-based cross-proteome comparative analyses.

Amino Acid Sequence↗

Data mining of sequences and 3D structures of allergenic proteins.

MOTIVATION: Many sequences, and in some cases structures, of proteins that induce an allergic response in atopic individuals have been determined in recent years. This data indicates that allergens, regardless of source, fall into discreet protein families. Similarities in the sequence may explain clinically observed cross-reactivities between different biological triggers. However, previously available allergy databases group allergens according to their biological sources, or observed clinical cross-reactivities, without providing data about the proteins. A computer-aided data mining system is needed to compare the sequential and structural details of known allergens. This information will aid in predicting allergenic cross-responses and eventually in determining possible common characteristics of IgE recognition. RESULTS: The new web-based Structural Database of Allergenic Proteins (SDAP) permits the user to quickly compare the sequence and structure of allergenic proteins. Data from literature sources and previously existing lists of allergens are combined in a MySQL interactive database with a wide selection of bioinformatics applications. SDAP can be used to rapidly determine the relationship between allergens and to screen novel proteins for the presence of IgE or T-cell epitopes they may share with known allergens. Further, our novel similarity search method, based on five dimensional descriptors of amino acid properties, can be used to scan the SDAP entries with a peptide sequence. For example, when a known IgE binding epitope from shrimp tropomyosin was used as a query, the method rapidly identified a similar sequence in known shellfish and insect allergens. This prediction of cross-reactivity between allergens is consistent with clinical observations. AVAILABILITY: SDAP is available on the web at http://fermi.utmb.edu/SDAP/index.html

Allergens↗

PSST-2.0: Protein Data Bank Sequence Search Tool.

UNLABELLED: PSST-2.0 (Protein Data Bank [PDB] Sequence Search Tool) is an updated version of the earlier PSST (Protein Sequence Search Tool), and the philosophy behind the search engine has remained unchanged. PSST-2.0 is a Web-based, interactive search engine developed to retrieve required protein or nucleic acid sequence information and some of its related details, primarily from sequences derived from the structures deposited in the PDB (the database of 3-dimensional [3-D] protein and nucleic acid structures). Additionally, the search engine works for a selected subset of 25% or 90% non-homologous protein chains. For some of the selected options, the search engine produces a detailed output for the user-uploaded, 3-D atomic coordinates of the protein structure (PDB file format) from the client machine through the Web browser. The search engine works on a locally maintained PDB, which is updated every week from the parent server at the Research Collaboratory for Structural Bioinformatics, and hence the search results are up to date at any given time. AVAILABILITY: PSST-2.0 is freely accessible via http://pranag.physics.iisc.ernet.in/psst/ or http://144.16.71.10/psst/.

Algorithms↗

Customization in a unified framework for summarizing medical literature.

OBJECTIVE: We present the summarization system in the PErsonalized Retrieval and Summarization of Images, Video and Language (PERSIVAL) medical digital library. Although we discuss the context of our summarization research within the PERSIVAL platform, the primary focus of this article is on strategies to define and generate customized summaries. METHODS AND MATERIAL: Our summarizer employs a unified user model to create a tailored summary of relevant documents for either a physician or lay person. The approach takes advantage of regularities in medical literature text structure and content to fulfill identified user needs. RESULTS: The resulting summaries combine both machine-generated text and extracted text that comes from multiple input documents. Customization includes both group-based modeling for two classes of users, physician and lay person, and individually driven models based on a patient record. CONCLUSIONS: Our research shows that customization is feasible in a medical digital library.

Abstracting and Indexing↗

MMDB: 3D structure data in Entrez.

Three-dimensional structures are now known for roughly half of all protein families. It is thus quite likely, in searching sequence databases, that one will encounter a homolog with known structure and be able to use this information to infer structure-function properties. The goal of Entrez's 3D structure database is to make this information accessible and useful to molecular biologists. To this end, Entrez's search engine provides three powerful features: (i) Links between databases; one may search by term matching in Medline((R)), for example, and link to 3D structures reported in these articles. (ii) Sequence and structure neighbors; one may select all sequences similar to one of interest, for example, and link to any known 3D structures. (iii) Sequence and structure visualization; identifying a homolog with known structure, one may view a combined molecular-graphic and alignment display, to infer approximate 3D structure. Entrez's MMDB (Molecular Modeling DataBase) may be accessed at: http://www.ncbi.nlm.nih.gov/Entrez/structure.html

Amino Acid Sequence↗

Automatic lexeme acquisition for a multilingual medical subword thesaurus.

PURPOSE: We present a method for the automated acquisition of a multilingual medical lexicon (for Spanish, French and Swedish) to be used within the framework of a medical cross-language text retrieval system. METHODS: For the lexical acquisition process, we incorporate seed lexicons and lists of trusted term translations derived from the UMLS Metathesaurus. The seed lexicons for Spanish, French and Swedish are automatically generated from (previously manually constructed) Portuguese, German and English sources by simple string transformations. Lexical and semantic hypotheses are then validated by processing pairs of term translations. In a last step, we use the cleaned list of "approved" translations in order to augment, step by step, the target dictionaries by processing the parallel corpora in terms of co-occurrence patterns of hypothesized translation equivalents which cannot be derived by simple character substitutions. RESULTS: An existing multilingual lexicon for the medical domain with about 60,000 entries for English, German, and Portuguese was automatically augmented by more then 17,000 new lexemes for Spanish, French, and Swedish. CONCLUSIONS: Our approach constitutes a promising method for the automated creation of new lexicon entries and their linkage to semantic identifiers.

Electronic Data Processing↗

Systematic literature review for clinical practice guideline development.

The purpose of this paper was to evaluate the quality and scope of the published literature on functional impairment due to cataract in adults as reviewed for the Agency for Health Care Policy and Research Clinical Practice Guideline. We examined the method of literature retrieved and analysis performed in the course of development of literature-based recommendations for the guideline panel. To collect data, we reviewed the process of literature acquisition and identification and the quality assessments made by reviewers of 14 individual topics composed of 77 issues related to the guideline. We collated this information to provide an assessment of the quality and scope of the relevant literature. Less than 4% (310) of the approximately 8,000 articles initially identified as potentially relevant to the guideline were ultimately used. The majority covered three topics (surgery and complication, 100; Nd:YAG capsulotomy, 77; and potential vision testing, 40). Three other topics--indications for surgery, preoperative medical evaluation, and rehabilitation--were devoid of articles meeting inclusion criteria. For 43 issues, there was no identifiable relevant literature. With few exceptions, the quality of the literature was rated fair to poor owing to major flaws in experimental design. Case series (256 reports) of one type or another accounted for the majority of the included literature. There were 17 random controlled trials. This review revealed a sparse and generally low-quality literature relevant to the management of functional impairment due to cataract, despite a relatively large data base in reputable peer-reviewed journals.

Adult↗

The tmRNA Website: invasion by an intron.

tmRNA (also known as 10Sa RNA or SsrA) plays a central role in an unusual mode of translation, whereby a stalled ribosome switches from a problematic mRNA to a short reading frame within tmRNA during translation of a single polypeptide chain. Research on the mechanism, structure and biology of tmRNA is served by the tmRNA Website, a collection of sequences for tmRNA and the encoded proteolysis-inducing peptide tags, alignments, careful documentation and other information; the URL is http://www.indiana.edu/~tmrna. Four pseudoknots are usually present in each tmRNA, so the database is rich with information on pseudoknot variability. Since last year it has doubled (227 tmRNA sequences as of September 2001), a sequence alignment for the tmRNA cofactor SmpB has been included, and genomic data for Clostridium botulinum has revealed a group I (subgroup IA3) intron interrupting the tmRNA T-loop.

Base Sequence↗

MannDB - a microbial database of automated protein sequence analyses and evidence integration for protein characterization.

BACKGROUND: MannDB was created to meet a need for rapid, comprehensive automated protein sequence analyses to support selection of proteins suitable as targets for driving the development of reagents for pathogen or protein toxin detection. Because a large number of open-source tools were needed, it was necessary to produce a software system to scale the computations for whole-proteome analysis. Thus, we built a fully automated system for executing software tools and for storage, integration, and display of automated protein sequence analysis and annotation data. DESCRIPTION: MannDB is a relational database that organizes data resulting from fully automated, high-throughput protein-sequence analyses using open-source tools. Types of analyses provided include predictions of cleavage, chemical properties, classification, features, functional assignment, post-translational modifications, motifs, antigenicity, and secondary structure. Proteomes (lists of hypothetical and known proteins) are downloaded and parsed from Genbank and then inserted into MannDB, and annotations from SwissProt are downloaded when identifiers are found in the Genbank entry or when identical sequences are identified. Currently 36 open-source tools are run against MannDB protein sequences either on local systems or by means of batch submission to external servers. In addition, BLAST against protein entries in MvirDB, our database of microbial virulence factors, is performed. A web client browser enables viewing of computational results and downloaded annotations, and a query tool enables structured and free-text search capabilities. When available, links to external databases, including MvirDB, are provided. MannDB contains whole-proteome analyses for at least one representative organism from each category of biological threat organism listed by APHIS, CDC, HHS, NIAID, USDA, USFDA, and WHO. CONCLUSION: MannDB comprises a large number of genomes and comprehensive protein sequence analyses representing organisms listed as high-priority agents on the websites of several governmental organizations concerned with bio-terrorism. MannDB provides the user with a BLAST interface for comparison of native and non-native sequences and a query tool for conveniently selecting proteins of interest. In addition, the user has access to a web-based browser that compiles comprehensive and extensive reports. Access to MannDB is freely available at http://manndb.llnl.gov/.

Algorithms↗

A study on the development of a filing system for funduscopic images with a personal computer for Twin AMHTS.

A filing system for ocular funduscopic image data was developed by using a personal computer for the Twin AMHTS. The development of the system was tried as one of the data transfer system including image data between two similar AMHTSs named the Twin AMHTS through the information network system. The filing system is capable of storing 26782 data of ophthalmoscopic pictures with a data compression mode by using a magneto-optical disk (MOD) whose storage capacity of both sides is 616 MB. It takes no long time for retrieval and display of the image data in the filing system. Good quality of compression and decompression obtained and reproducibility of the ocular fundus picture is favorable regardless of normal or abnormal cases. As a result, it is suggested that the developed system has practical utility although it requires more improvement.

Computer Communication Networks↗