Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Storage And Retrieval”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

ALTER: eclectic management of molecular structure data.

ALTER is a computer program written to facilitate easy conversion between different representations of molecular structure data. The program functions as a file converter, data generation engine, and through the creation of control or input files, as an interface to other programs. The main aspects of program function--the reading and writing of files; coordinate transformation; data reorganization: structure building; data abstraction, including the generation of a wide variety of topological indices and constitutional descriptors; and display--are described in appropriate detail.

Computer Graphics↗

Characterization of the phosphorylation sites of human high molecular weight neurofilament protein by electrospray ionization tandem mass spectrometry and database searching.

Hyperphosphorylated high molecular weight neurofilament protein (NF-H) exhibits extensive phosphorylation on lysine-serine-proline (KSP) repeats in the C-terminal domain of the molecule. Specific phosphorylation sites in human NF-H were identified by proteolytic digestion and analysis of the resulting digests by a combination of microbore liquid chromatography, electrospray ionization tandem (MS/MS) ion trap mass spectrometry, and database searching. The computer programs utilized (PEPSEARCH and SEQUEST) are capable of identifying peptides and phosphorylation sites from uninterpreted MS/MS spectra, and by use of these methods, 27 phosphopeptides and their phosphorylated residues were identified. On the basis of these phosphopeptides, 38 phosphorylation sites in human NF-H were characterized. These include 33 KSP, lysine-threonine-proline (KTP) or arginine-serine-proline (RSP) sites and four unphosphorylated sites, all of which occur in the KSP repeat domain (residues 502-823); and one threonine phosphorylation site observed in a KVPTPEK motif. Six KSP sites were not characterized because of the failure to isolate and identify corresponding phosphopeptides. Heterogeneity in serine and threonine phosphorylation was observed at three sites or deduced to occur at three sites on the basis of enzyme specificity. As a result of the phosphorylated motifs identified (KSPAKEE, KSPVKEE, KS/TPEKAK, KSPEKEE, KSPVKAE, KSPAEAK, KSPPEAK, KSPEAKT, KSPAEVK, and KVPTPEK), human NF-H tail domain is postulated to be a substrate of proline-directed kinases. The threonine-phosphorylated KVPTPEK motif suggested the existence of a novel proline-directed kinase.

Amino Acid Sequence↗

Medium scale integration of molecular logic gates in an automaton.

The assembly of molecular automata that perform increasingly complex tasks, such as game playing, presents an unbiased test of molecular computation. We now report a second-generation deoxyribozyme-based automaton, MAYA-II, which plays a complete game of tic-tac-toe according to a perfect strategy. In silicon terminology, MAYA-II represents the first "medium-scale integrated molecular circuit", integrating 128 deoxyribozyme-based logic gates, 32 input DNA molecules, and 8 two-channel fluorescent outputs across 8 wells.

Algorithms↗

BioMagResBank database with sets of experimental NMR constraints corresponding to the structures of over 1400 biomolecules deposited in the Protein Data Bank.

Experimental constraints associated with NMR structures are available from the Protein Data Bank (PDB) in the form of "Magnetic Resonance" (MR) files. These files contain multiple types of data concatenated without boundary markers and are difficult to use for further research. Reported here are the results of a project initiated to annotate, archive, and disseminate these data to the research community from a searchable resource in a uniform format. The MR files from a set of 1410 NMR structures were analyzed and their original constituent data blocks annotated as to data type using a semi-automated protocol. A new software program called Wattos was then used to parse and archive the data in a relational database. From the total number of MR file blocks annotated as constraints, it proved possible to parse 84% (3337/3975). The constraint lists that were parsed correspond to three data types (2511 distance, 788 dihedral angle, and 38 residual dipolar couplings lists) from the three most popular software packages used in NMR structure determination: XPLOR/CNS (2520 lists), DISCOVER (412 lists), and DYANA/DIANA (405 lists). These constraints were then mapped to a developmental version of the BioMagResBank (BMRB) data model. A total of 31 data types originating from 16 programs have been classified, with the NOE distance constraint being the most commonly observed. The results serve as a model for the development of standards for NMR constraint deposition in computer-readable form. The constraints are updated regularly and are available from the BMRB web site (http://www.bmrb.wisc.edu).

Databases, Protein↗

Issues in searching molecular sequence databases.

Sequence similarity search programs are versatile tools for the molecular biologist, frequently able to identify possible DNA coding regions and to provide clues to gene and protein structure and function. While much attention had been paid to the precise algorithms these programs employ and to their relative speeds, there is a constellation of associated issues that are equally important to realize the full potential of these methods. Here, we consider a number of these issues, including the choice of scoring systems, the statistical significance of alignments, the masking of uninformative or potentially confounding sequence regions, the nature and extent of sequence redundancy in the databases and network access to similarity search services.

Algorithms↗

Xlandscape: the graphical display of word frequencies in sequences.

MOTIVATION: To provide a graphical interface for the generation, display and manipulation of a sequence landscape that will run on all X-windows-based Unix workstations. RESULTS: The sequence landscape approach enables the representation of the frequency of occurrence of all query sequence sub-words within a database. The landscape approach can detect tandem and other repeating word motifs, specific sub-words that are over-represented words in a particular database using Markov probability and the preference for sub-words belonging to either one of two databases. All these features aid in the classification of a query sequence. Given the open-text format for sequences and databases, the Xlandscape tool can be applied to a wide range of problems.

Algorithms↗

WebPHYLIP: a web interface to PHYLIP.

A web interface to PHYLIP (version 3.57 C) is implemented using CGI/Perl programming. It enables users to do phylogenetic analysis through the Internet.

Data Display↗

Automated extraction of information on protein-protein interactions from the biological literature.

MOTIVATION: To understand biological process, we must clarify how proteins interact with each other. However, since information about protein-protein interactions still exists primarily in the scientific literature, it is not accessible in a computer-readable format. Efficient processing of large amounts of interactions therefore needs an intelligent information extraction method. Our aim is to develop an efficient method for extracting information on protein-protein interaction from scientific literature. RESULTS: We present a method for extracting information on protein-protein interactions from the scientific literature. This method, which employs only a protein name dictionary, surface clues on word patterns and simple part-of-speech rules, achieved high recall and precision rates for yeast (recall = 86.8% and precision = 94.3%) and Escherichia coli (recall = 82.5% and precision = 93.5%). The result of extraction suggests that our method should be applicable to any species for which a protein name dictionary is constructed. AVAILABILITY: The program is available on request from the authors.

Electronic Data Processing↗

RED: the analysis, management and dissemination of expressed sequence tags.

The Rancourt EST Database (RED) is a web-based system for the analysis, management, and dissemination of expressed sequence tags (ESTs). RED represents a flexible template DNA sequence database that can be easily manipulated to suit the needs of other laboratories undertaking mid-size sequencing projects.

Animals↗

OWEN: aligning long collinear regions of genomes.

OWEN is an interactive tool for aligning two long DNA sequences that represents similarity between them by a chain of collinear local similarities. OWEN employs several methods for constructing and editing local similarities and for resolving conflicts between them. Alignments of sequences of lengths over 10(6) can often be produced in minutes. OWEN requires memory below 20 L, where L is the sum of lengths of the compared sequences.

Algorithms↗

TranScout: prediction of gene expression regulatory proteins from their sequences.

MOTIVATION: The advent of genomics yields thousands of reading frames in search of function. Identification of conserved functional motifs in protein sequences can be helpful for function prediction. RESULTS: A database and a classification of reported DNA-binding protein motifs has been designed. A program ('TranScout') has been developed for the detection and evaluation of conserved motifs in prokaryotic and eukaryotic sequences of proteins with a gene regulatory function. The efficiency of the program is shown in a benchmark against a database obtained from SWISS-PROT without the protein sequences used to train the program. All motifs were detected with a mean average sensitivity of 0.98 and a mean average specificity of 0.92. AVAILABILITY: The program is freely available for use on the internet at http://luz.uab.es/transcout/. The user can find additional information at this site.

Algorithms↗

Representation of DNA sequences with virtual potentials and their processing by (SEQREP) Kohonen self-organizing maps.

MOTIVATION: We propose representing individual positions in DNA sequences by virtual potentials generated by other bases of the same sequence. This is a compact representation of the neighbourhood of a base. The distribution of the virtual potentials over the whole sequence can be used as a representation of the entire sequence (SEQREP code). It is a flexible code, with a length independent of the sequence size, does not require previous alignment, and is convenient for processing by neural networks or statistical techniques. RESULTS: To evaluate its biological significance, the SEQREP code was used for training Kohonen self-organizing maps (SOMs) in two applications: (a) detection of Alu sequences, and (b) classification of sequences encoding for HIV-1 envelope glycoprotein (env) into subtypes A-G. It was demonstrated that SOMs clustered sequences belonging to different classes into distinct regions. For independent test sets, very high rates of correct predictions were obtained (97% in the first application, 91% in the second). Possible areas of application of SEQREP codes include functional genomics, phylogenetic analysis, detection of repetitions, database retrieval, and automatic alignment. AVAILABILITY: Software for representing sequences by SEQREP code, and for training Kohonen SOMs is made freely available from http://www.dq.fct.unl.pt/qoa/jas/seqrep. SUPPLEMENTARY INFORMATION: Supplementary material is available at http://www.dq.fct.unl.pt/qoa/jas/seqrep/bioinf2002

Algorithms↗

RnaViz 2: an improved representation of RNA secondary structure.

SUMMARY: RnaViz has been developed to easily create nice, publication quality drawings of RNA secondary structure. RnaViz 2 supports CT, DCSE, and RNAML input formats and improves on many aspects of the first version, notably portability and structure annotation. RnaViz is written using a hybrid programming approach combining pieces written in C and in the scripting language Tcl/Tk, making the program very portable and extensible. AVAILABILITY: Source code, binaries for Linux and MS Windows, and additional documentation are available athttp://rrna.uia.ac.be/rnaviz/

Base Sequence↗

The Z curve database: a graphic representation of genome sequences.

MOTIVATION: Genome projects for many prokaryotic and eukaryotic species have been completed and more new genome projects are being underway currently. The availability of a large number of genomic sequences for researchers creates a need to find graphic tools to study genomes in a perceivable form. The Z curve is one of such tools available for visualizing genomes. The Z curve is a unique three-dimensional curve representation for a given DNA sequence in the sense that each can be uniquely reconstructed given the other. The Z curve database for more than 1000 genomes have been established here. RESULTS: The database contains the Z curves for archaea, bacteria, eukaryota, organelles, phages, plasmids, viroids and viruses, whose genomic sequences are currently available. All the 3-dimensional Z curves and their three component curves are stored in the database. The applications of the Z curve database on comparative genomics, gene prediction, computation of G+C content with a windowless technique, prediction of replication origins and terminations of bacterial and archaeal genomes and study of local deviations from the Chargaff Parity Rule 2 etc. are presented in detail. The Z curve database reported here is a treasure trove in which biologists could find useful biological knowledge.

Animals↗