Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Storage And Retrieval”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

A new family of carbon-nitrogen hydrolases.

Using computer methods for database search and multiple alignment, statistically significant sequence similarities were identified between several nitrilases with distinct substrate specificity, cyanide hydratases, aliphatic amidases, beta-alanine synthase, and a few other proteins with unknown molecular function. All these proteins appear to be involved in the reduction of organic nitrogen compounds and ammonia production. Sequence conservation over the entire length, as well as the similarity in the reactions catalyzed by the known enzymes in this family, points to a common catalytic mechanism. The new family of enzymes is characterized by several conserved motifs, one of which contains an invariant cysteine that is part of the catalytic site in nitrilases. Another highly conserved motif includes an invariant glutamic acid that might also be involved in catalysis.

Amidohydrolases↗

Fast structure alignment for protein databank searching.

A fast method is described for searching and analyzing the protein structure databank. It uses secondary structure followed by residue matching to compare protein structures and is developed from a previous structural alignment method based on dynamic programming. Linear representations of secondary structures are derived and their features compared to identify equivalent elements in two proteins. The secondary structure alignment then constrains the residue alignment, which compares only residues within aligned secondary structures and with similar buried areas and torsional angles. The initial secondary structure alignment improves accuracy and provides a means of filtering out unrelated proteins before the slower residue alignment stage. It is possible to search or sort the protein structure databank very quickly using just secondary structure comparisons. A search through 720 structures with a probe protein of 10 secondary structures required 1.7 CPU hours on a Sun 4/280. Alternatively, combined secondary structure and residue alignments, with a cutoff on the secondary structure score to remove pairs of unrelated proteins from further analysis, took 10.1 CPU hours. The method was applied in searches on different classes of proteins and to cluster a subset of the databank into structurally related groups. Relationships were consistent with known families of protein structure.

Amino Acid Sequence↗

MS1, MS2, and SQT-three unified, compact, and easily parsed file formats for the storage of shotgun proteomic spectra and identifications.

As the speed with which proteomic labs generate data increases along with the scale of projects they are undertaking, the resulting data storage and data processing problems will continue to challenge computational resources. This is especially true for shotgun proteomic techniques that can generate tens of thousands of spectra per instrument each day. One design factor leading to many of these problems is caused by storing spectra and the database identifications for a given spectrum as individual files. While these problems can be addressed by storing all of the spectra and search results in large relational databases, the infrastructure to implement such a strategy can be beyond the means of academic labs. We report here a series of unified text file formats for storing spectral data (MS1 and MS2) and search results (SQT) that are compact, easily parsed by both machine and humans, and yet flexible enough to be coupled with new algorithms and data-mining strategies.

Database Management Systems↗

Detection of signature sequences in overlapping genes and prediction of a novel overlapping gene in hepatitis G virus.

In viruses an increased coding ability is provided by overlapping genes, in which two alternative open reading frames (ORFs) may be translated to yield two distinct proteins. The identification of signature sequences in overlapping genes is a topic of particular interest, since additional out-of-frame coding regions can be nested within known genes. In this work, a novel feature peculiar to overlapping coding regions is presented. It was detected by analysis of a sample set of 21 virus genomic sequences and consisted in the repeated occurrence of a cluster of basic amino acid residues, encoded by a frame, combined to a stretch of acidic residues, encoded by the corresponding overlapping frame. A computer scan of an additional set of virus sequences demonstrated that this feature is common to several other known overlapping ORFs and led to prediction of a novel overlapping gene in hepatitis G virus (HGV). The occurrence of a bifunctional coding region in HGV was also supported by its extremely lower rate of synonymous nucleotide substitutions compared to that observed in the other gene regions of the HGV genome. Analysis of the amino acid sequence that was deduced from the putative overlapping gene revealed a high content of basic residues and the presence of a nuclear targeting signal; these characteristics suggest that a core-like protein may be expressed by this novel ORF.

Algorithms↗

Real-time remote telefluoroscopic assessment of patients with dysphagia.

Dysphagia is a serious health problem that affects persons of all ages, from the neonate to those of advanced age. Many smaller communities and areas with sparse populations do not have regular access to professionals with expertise in the area of oral/pharyngeal dysphagia. Telemedicine is one method by which people in these areas can receive quality of service. The intent of this work was to develop an Internet system that permits real-time, remote, interactive evaluation of oral/pharyngeal swallowing function. The system consists of two major components. The first is a PC that is located in the fluoroscopy suite of a hospital. The computer is connected to the fluoroscope output and is responsible for (1) capturing video signals, (2) converting the analog video data into digitized video formats of both full resolution and transmission-optimized resolution, (3) simultaneously transmitting the transmission-optimized video stream over the network while the examination is being performed, and (4) storing the full-resolution data as a file in local storage for later retrieval. The second component is the controller computer which is located at a site some distance from the hospital. That controller computer manages the video capture process at the remote hospital site, manages the transmission of the stored images, and is then used for video analysis. The delay between the image as it was captured at the remote hospital site and viewed on the controller computer in the Principal Investigator's (PI's) laboratory ranged from 3 to 5 s. Video transmission occurred over a standard T1 line.

Computer Communication Networks↗

MULTI: a shared memory approach to cooperative molecular modeling.

A general purpose molecular modeling system, MULTI, based on the UNIX shared memory and semaphore facilities for interprocess communication is described. In addition to the normal querying or monitoring of geometric data, MULTI also provides processes for manipulating conformations, and for displaying peptide or nucleic acid ribbons, Connolly surfaces, close nonbonded contacts, crystal-symmetry related images, least-squares superpositions, and so forth. This paper outlines the basic techniques used in MULTI to ensure cooperation among these specialized processes, and then describes how they can work together to provide a flexible modeling environment.

Computer Communication Networks↗

Automated construction and graphical presentation of protein blocks from unaligned sequences.

Protein blocks consist of multiply aligned sequence segments that correspond to the most highly conserved regions of protein families. Typically, a set of related proteins has more than one region in common and their relationship can be represented as a series of ungapped blocks separated by unaligned regions. Blockmaker is an automated system available by electronic mail (blockmaker@howard.fhcrc.org) and the World Wide Web (http://www.blocks.fhcrc.org4) that finds blocks in a group of related protein sequences submitted by the user. It adapts and extends existing algorithms to make them useful to biologists looking for conserved regions in a group of related proteins sequences. Two sets of blocks are returned, one in which candidate blocks are detected using the MOTIF algorithm and the other using a Gibbs sampler algorithm that has been adapted for full automation. This use of two block-finding methods based on completely different principles provides a 'reality check,' whereby a block detected by both methods is considered to be correct. Resulting blocks can be displayed using the information-based 'sequence logo' method, adapted to incorporate sequence weights, which provides an intuitive visual description of both the residue and the conservation information at each position. Blocks generated by this system are useful in diverse applications, such as searching databases and designing degenerate PCR primers. As an example, blocks made from amino acid sequences related to Caenorhabditis elegans Tc1 transposase were used to search GenBank, revealing that several fish and amphibian genomic sequences harbor previously unreported Tc1 homologs.

Algorithms↗

Searching protein sequence libraries: comparison of the sensitivity and selectivity of the Smith-Waterman and FASTA algorithms.

The sensitivity and selectivity of the FASTA and the Smith-Waterman protein sequence comparison algorithms were evaluated using the superfamily classification provided in the National Biomedical Research Foundation/Protein Identification Resource (PIR) protein sequence database. Sequences from each of the 34 superfamilies in the PIR database with 20 or more members were compared against the protein sequence database. The similarity scores of the related and unrelated sequences were determined using either the FASTA program or the Smith-Waterman local similarity algorithm. These two sets of similarity scores were used to evaluate the ability of the two comparison algorithms to identify distantly related protein sequences. The FASTA program using the ktup = 2 sensitivity setting performed as well as the Smith-Waterman algorithm for 19 of the 34 superfamilies. Increasing the sensitivity by setting ktup = 1 allowed FASTA to perform as well as Smith-Waterman on an additional 7 superfamilies. The rigorous Smith-Waterman method performed better than FASTA with ktup = 1 on 8 superfamilies, including the globins, immunoglobulin variable regions, calmodulins, and plastocyanins. Several strategies for improving the sensitivity of FASTA were examined. The greatest improvement in sensitivity was achieved by optimizing a band around the best initial region found for every library sequence. For every superfamily except the globins and immunoglobulin variable regions, this strategy was as sensitive as a full Smith-Waterman. For some sequences, additional sensitivity was achieved by including conserved but nonidentical residues in the lookup table used to identify the initial region.

Algorithms↗

DNA workbench.

Explore the source record for details and available documents.

Base Sequence↗

Temporal abstraction in intelligent clinical data analysis: a survey.

OBJECTIVE: Intelligent clinical data analysis systems require precise qualitative descriptions of data to enable effective and context sensitive interpretation to take place. Temporal abstraction (TA) provides the means to achieve such descriptions, which can then be used as input to a reasoning engine where they are evaluated against a knowledge base to arrive at possible clinical hypotheses. This paper surveys previous research into the development of intelligent clinical data analysis systems that incorporate TA mechanisms and presents research synergies and trends across the research reviewed, especially those associated with the multi-dimensional nature of real-time patient data streams. The motivation for this survey is case study based research into the development of an intelligent real-time, high-frequency patient monitoring system to provide detection of temporal patterns within multiple patient data streams. RESULTS: The survey was based on factors that are of importance to broaden research into temporal abstraction and on characteristics we believe will assume an increasing level of importance for future clinical IDA systems. These factors were: aspects of the data that is abstracted such as source domain and sample frequency, complexity available within abstracted patterns, dimensionality of the TA and data environment and the knowledge and reasoning underpinning TA processes. CONCLUSION: It is evident from the review that for intelligent clinical data analysis systems to progress into the future where clinical environments are becoming increasingly data-intensive, the ability for managing multi-dimensional aspects of data at high observation and sample frequencies must be provided. Also, the detection of complex patterns within patient data requires higher levels of TA than are presently available. The conflicting matters of computational tractability and temporal reasoning within a real-time environment present a non-trivial problem for investigation in regard to these matters. Finally, to be able to fully exploit the value of learning new knowledge from stored clinical data through data mining and enable its application to data abstraction, the fusion of data mining and TA processes becomes a necessity.

Artificial Intelligence↗

Prediction of protein subcellular locations by GO-FunD-PseAA predictor.

The localization of a protein in a cell is closely correlated with its biological function. With the explosion of protein sequences entering into DataBanks, it is highly desired to develop an automated method that can fast identify their subcellular location. This will expedite the annotation process, providing timely useful information for both basic research and industrial application. In view of this, a powerful predictor has been developed by hybridizing the gene ontology approach [Nat. Genet. 25 (2000) 25], functional domain composition approach [J. Biol. Chem. 277 (2002) 45765], and the pseudo-amino acid composition approach [Proteins Struct. Funct. Genet. 43 (2001) 246; Erratum: ibid. 44 (2001) 60]. As a showcase, the recently constructed dataset [Bioinformatics 19 (2003) 1656] was used for demonstration. The dataset contains 7589 proteins classified into 12 subcellular locations: chloroplast, cytoplasmic, cytoskeleton, endoplasmic reticulum, extracellular, Golgi apparatus, lysosomal, mitochondrial, nuclear, peroxisomal, plasma membrane, and vacuolar. The overall success rate of prediction obtained by the jackknife cross-validation was 92%. This is so far the highest success rate performed on this dataset by following an objective and rigorous cross-validation procedure.

Algorithms↗

Hum-PLoc: a novel ensemble classifier for predicting human protein subcellular localization.

Predicting subcellular localization of human proteins is a challenging problem, especially when unknown query proteins do not have significant homology to proteins of known subcellular locations and when more locations need to be covered. To tackle the challenge, protein samples are expressed by hybridizing the gene ontology (GO) database and amphiphilic pseudo amino acid composition (PseAA). Based on such a representation frame, a novel ensemble classifier, called "Hum-PLoc", was developed by fusing many basic individual classifiers through a voting system. The "engine" of these basic classifiers was operated by the KNN (K-nearest neighbor) rule. As a demonstration, tests were performed with the ensemble classifier for human proteins among the following 12 locations: (1) centriole; (2) cytoplasm; (3) cytoskeleton; (4) endoplasmic reticulum; (5) extracell; (6) Golgi apparatus; (7) lysosome; (8) microsome; (9) mitochondrion; (10) nucleus; (11) peroxisome; (12) plasma membrane. To get rid of redundancy and homology bias, none of the proteins investigated here had > or = 25% sequence identity to any other in a same subcellular location. The overall success rates thus obtained via the jackknife cross-validation test and independent dataset test were 81.1% and 85.0%, respectively, which are more than 50% higher than those obtained by the other existing methods on the same stringent datasets. Furthermore, an incisive and compelling analysis was given to elucidate that the overwhelmingly high success rate obtained by the new predictor is by no means due to a trivial utilization of the GO annotations. This is because, for those proteins with "subcellular location unknown" annotation in Swiss-Prot database, most (more than 99%) of their corresponding GO numbers in GO database are also annotated with "cellular component unknown". The information and clues for predicting subcellular locations of proteins are actually buried into a series of tedious GO numbers, just like they are buried into a pile of complicated amino acid sequences although with a different manner and "depth". To dig out the knowledge about their locations, a sophisticated operation engine is needed. And the current predictor is one of these kinds, and has proved to be a very powerful one. The Hum-PLoc classifier is available as a web-server at http://202.120.37.186/bioinf/hum.

Algorithms↗

A protein database constructed from low-coverage genomic sequence of Bacillus megaterium and its use for accelerated proteomic analysis.

Peptide mass fingerprint (PMF) matching is a high-throughput method used for protein spot identification in connection with two-dimensional gel electrophoresis (2DE). However, the success of PMF matching largely depends on whether the proteins to be identified exist in the database searched. Consequently, it is often necessary to apply other more sophisticated but also time-consuming technologies to generate sequence-tags for definitive protein identification. On the other hand, modern sequencing technologies are generating a large quantity of DNA sequences, first in unfinished form or with low genome coverage due to the time-consuming and thus limiting steps of finishing and annotation. We recently started to sequence the genome of Bacillus megaterium DSM 319, a bacterium of industrial interest. In this study, we demonstrate that a protein database generated from merely three-fold coverage, unfinished genomic sequences of this bacterium allows a fast and reliable protein spot identification solely based on PMF from high-throughput MALDI-TOF MS analysis. We further show that the strain-specific protein database from low coverage genomic sequence greatly outperforms the commonly used cross-species databases constructed from 13 completely sequenced Bacillus strains for protein spot identification via PMF.

Algorithms↗

Towards a formalization of disease-specific ontologies for neuroinformatics.

We present issues arising when trying to formalize disease maps, i.e. ontologies to represent the terminological relationships among concepts necessary to construct a knowledge-base of neurological disorders. These disease maps are being created in the context of a large-scale data mediation system being created for the Biomedical Informatics Research Network (BIRN). The BIRN is a multi-university consortium collaborating to establish a large-scale data and computational grid around neuroimaging data, collected across multiple scales. Test bed projects within BIRN involve both animal and human studies of Alzheimer's disease, Parkinson's disease and schizophrenia. Incorporating both the static 'terminological' relationships and dynamic processes, disease maps are being created to encapsulate a comprehensive theory of a disease. Terms within the disease map can also be connected to the relevant terms within other ontologies (e.g. the Unified Medical Language System), in order to allow the disease map management system to derive relationships between a larger set of terms than what is contained within the disease map itself. In this paper, we use the basic structure of a disease map we are developing for Parkinson's disease to illustrate our initial formalization for disease maps.

Electronic Data Processing↗

Representation and searching of carbohydrate structures using graph-theoretic techniques.

This paper describes how the carbohydrate structures in the Complex Carbohydrate Structure Database (CCSD) can be represented by labelled graphs, in which the nodes and edges of a graph are used to denote the residues and the inter-residue linkages, respectively, of a carbohydrate. These graph representations are then searched with a subgraph-isomorphism algorithm. We describe the use of one such algorithm, that due to Ullmann, and demonstrate that it provides a very precise way of searching the structures in CCSD. We also describe the use of screening techniques that can eliminate many of the CCSD structures from the subgraph-isomorphism search, with a consequent increase in the speed of the search.

Algorithms↗

Finding the hairpin in the haystack: searching for RNA motifs.

A growing list of examples underscores the roles that regulatory RNA motifs play in controlling the genetic repertoire of cells and developing organisms. Once either an RNA-processing signal, a ribozyme, an element that controls translational or mRNA stability or an RNA localization signal has been identified, it is important to search for other RNA sequences that bear similar regulatory signals. While DNA regulatory elements can often be described by a consensus sequence, RNA signals are frequently composed of a combination of sequence and structure motifs. Here, we discuss the approaches that can be used to identify RNA motifs by searching databases.

Animals↗

Using BLAST for identifying gene and protein names in journal articles.

We describe a system which automatically identifies gene and protein names in journal articles, an important and non-trivial first step in knowledge extraction of protein and gene actions. Our system uses a database of gene and protein names and is based on BLAST [Altschul et al., Nucleic Acids Res. 25 (1997) 3389-3402], a popular tool for DNA and protein sequence comparison. We describe a method that consists of mapping sequences of text characters into sequences of nucleotides that can be processed by BLAST. We demonstrate that this approach is feasible: the system matches gene and protein names with a recall of 78.8% and a precision of 71.7%, which includes names that are not part of the system database. An analysis of the results suggests techniques that can be used to improve performance further.

Algorithms↗