Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Storage And Retrieval”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 829 records · Page 46Linked to original sources

Rapid 3D protein structure database searching using information retrieval techniques.

MOTIVATION: As the sizes of three-dimensional (3D) protein structure databases are growing rapidly nowadays, exhaustive database searching, in which a 3D query structure is compared to each and every structure in the database, becomes inefficient. We propose a rapid 3D protein structure retrieval system named 'ProtDex2', in which we adopt the techniques used in information retrieval systems in order to perform rapid database searching without having access to every 3D structure in the database. The retrieval process is based on the inverted-file index constructed on the feature vectors of the relationships between the secondary structure elements (SSEs) of all the 3D protein structures in the database. ProtDex2 is a significant improvement, both in terms of speed and accuracy, upon its predecessor system, ProtDex. RESULTS: The experimental results show that ProtDex2 is very much faster than two well-known protein structure comparison methods, DALI and CE, yet not sacrificing on the accuracy of the comparison. When comparing with a similar SSE-based method, namely TopScan, ProtDex2 is much faster with comparable degree of accuracy. AVAILABILITY: The software is available at: http://xena1.ddns.comp.nus.edu.sg/~genesis/PD2.htm

Algorithms↗

YAdumper: extracting and translating large information volumes from relational databases to structured flat files.

Downloading the information stored in relational databases into XML and other flat formats is a common task in bioinformatics. This periodical dumping of information requires considerable CPU time, disk and memory resources. YAdumper has been developed as a purpose-specific tool to deal with the integral structured information download of relational databases. YAdumper is a Java application that organizes database extraction following an XML template based on an external Document Type Declaration. Compared with other non-native alternatives, YAdumper substantially reduces memory requirements and considerably improves writing performance.

Algorithms↗

Hierarchical classification of hydrolases catalytic sites.

UNLABELLED: Universal ontology of catalytic sites is required to systematize enzyme catalytic sites, their evolution as well as relations between catalytic sites and protein families, organisms and chemical reactions. Here we present a classification of hydrolases catalytic sites based on hierarchical organization. The web-accessible database provides information on the catalytic sites, protein folds, EC numbers and source organisms of the enzymes and includes software allowing for analysis and visualization of the relations between them. AVAILABILITY: http://www.enzyme.chem.msu.ru/hcs/

Amino Acid Sequence↗

Identification of post-translational modifications via blind search of mass-spectra.

Post-translational modifications (PTMs) are of great biological importance. Most existing approaches perform a restrictive search that can only take into account a few types of PTMs and ignore all others. We describe an unrestrictive PTM search algorithm that searches for all types of PTMs at once in a blind mode, i.e., without knowing which PTMs exist in a sample. The blind PTM identification opens a possibility to study the extent and frequencies of different types of PTMs, still an open problem in proteomics. Using our new algorithm, we were able to construct a two-dimensional PTM frequency matrix that reflects the number of MS/MS spectra in a sample for each putative PTM type and each amino acid. Application of this approach to a large IKKb dataset resulted in the largest set of PTMs reported for a single MS/MS sample so far. We demonstrate an excellent correlation between high values in the PTM frequency matrix and known PTMs thus validating our approach. We further argue that the PTM frequency matrix may reveal some still unknown modifications that warrant further experimental validation.

Algorithms↗

A technique for producing scalable color-quantized images with error diffusion.

To reliably and efficiently deliver media information to diverse clients over heterogeneous networks, the media involved must be scalable. In this paper, a color quantization algorithm for generating scalable color-indexed images is proposed based on a multiscale error diffusion framework. Images of lower resolutions are embedded in the outputs such that a simple down-sampling process can extract images of any desirable resolutions. Images possessing this scalable property support transmission over the Internet which contains clients with different display resolutions, systems with different caching resources and networks with varying bandwidths and QoS capabilities. Unlike most of the color halftoning algorithms available nowadays, the proposed algorithm is not dedicated for printing applications but for color-indexed displays. It works with any arbitrary palettes of different size.

Algorithms↗

Layered Wyner-Ziv video coding.

Following recent theoretical works on successive Wyner-Ziv coding (WZC), we propose a practical layered Wyner-Ziv video coder using the DCT, nested scalar quantization, and irregular LDPC code based Slepian-Wolf coding (or lossless source coding with side information at the decoder). Our main novelty is to use the base layer of a standard scalable video coder (e.g., MPEG-4/H.26L FGS or H.263+) as the decoder side information and perform layered WZC for quality enhancement. Similar to FGS coding, there is no performance difference between layered and monolithic WZC when the enhancement bitstream is generated in our proposed coder. Using an H.26L coded version as the base layer, experiments indicate that WZC gives slightly worse performance than FGS coding when the channel (for both the base and enhancement layers) is noiseless. However, when the channel is noisy, extensive simulations of video transmission over wireless networks conforming to the CDMA2000 1X standard show that H.26L base layer coding plus Wyner-Ziv enhancement layer coding are more robust against channel errors than H.26L FGS coding. These results demonstrate that layered Wyner-Ziv video coding is a promising new technique for video streaming over wireless networks.

Algorithms↗

Long-term forecasting of internet backbone traffic.

We introduce a methodology to predict when and where link additions/upgrades have to take place in an Internet protocol (IP) backbone network. Using simple network management protocol (SNMP) statistics, collected continuously since 1999, we compute aggregate demand between any two adjacent points of presence (PoPs) and look at its evolution at time scales larger than 1 h. We show that IP backbone traffic exhibits visible long term trends, strong periodicities, and variability at multiple time scales. Our methodology relies on the wavelet multiresolution analysis (MRA) and linear time series models. Using wavelet MRA, we smooth the collected measurements until we identify the overall long-term trend. The fluctuations around the obtained trend are further analyzed at multiple time scales. We show that the largest amount of variability in the original signal is due to its fluctuations at the 12-h time scale. We model inter-PoP aggregate demand as a multiple linear regression model, consisting of the two identified components. We show that this model accounts for 98% of the total energy in the original signal, while explaining 90% of its variance. Weekly approximations of those components can be accurately modeled with low-order autoregressive integrated moving average (ARIMA) models. We show that forecasting the long term trend and the fluctuations of the traffic at the 12-h time scale yields accurate estimates for at least 6 months in the future.

Algorithms↗

The perceptual scalability of visualization.

Larger, higher resolution displays can be used to increase the scalability of information visualizations. But just how much can scalability increase using larger displays before hitting human perceptual or cognitive limits? Are the same visualization techniques that are good on a single monitor also the techniques that are best when they are scaled up using large, high-resolution displays? To answer these questions we performed a controlled experiment on user performance time, accuracy, and subjective workload when scaling up data quantity with different space-time-attribute visualizations using a large, tiled display. Twelve college students used small multiples, embedded bar matrices, and embedded time-series graphs either on a 2 megapixel (Mp) display or with data scaled up using a 32 Mp tiled display. Participants performed various overview and detail tasks on geospatially-referenced multidimensional time-series data. Results showed that current designs are perceptually scalable because they result in a decrease in task completion time when normalized per number of data attributes along with no decrease in accuracy. It appears that, for the visualizations selected for this study, the relative comparison between designs is generally consistent between display sizes. However, results also suggest that encoding is more important on a smaller display while spatial grouping is more important on a larger display. Some suggestions for designers are provided based on our experience designing visualizations for large displays.

Computer Graphics↗

Prospects for single-molecule information-processing devices for the next paradigm.

Present information technologies use semiconductor devices and magnetic/optical disks; however, they are all foreseen to face fundamental limitations within a decade. Therefore, superceding devices are required for the next paradigm of high-performance information technologies. The paper first describes architectures suitable for single-molecule information processing, in which it is claimed that the performance of information processing is higher if speed and element number product is larger in almost all known architectures. Thus, single-molecule information-processing devices should be the most appropriate approach for the next paradigm. Then, prospects for single-molecule devices suitable for future information-processing technologies are described. Four possible milestones for realizing the Peta/Exa-floating operations per second (FLOPS) personal molecular supercomputer are proposed. Current status and necessary technologies of the first milestone are described, and necessary technologies for the next three milestones are also discussed.

Computers↗

Comparison of information processing technologies.

OBJECTIVE: To examine the type of information obtainable from scientific papers, using three different methods for the extraction, organization, and preparation of literature reviews. DESIGN: A set of three review papers was identified, and the ideas represented by the authors of those papers were extracted. The 161 articles referenced in those three reviews were then analyzed using 1) a formalized data extraction approach, which uses a protocol-driven manual process to extract the variables, values, and statistical significance of the stated relationships; and 2) a computerized approach known as "Idea Analysis," which uses the abstracts of the original articles and processes them through a computer software program that reads the abstracts and organizes the ideas presented by the authors. The results were then compared. The literature focused on the human papillomavirus and its relationship to cervical cancer. RESULTS: Idea Analysis was able to identify 68.9 percent of the ideas considered by the authors of the three review papers to be of importance in describing the association between human papillomavirus and cervical cancer. The formalized data extraction identified 27 percent of the authors' ideas. The combination of the two approaches identified 74.3 percent of the ideas considered important in the relationship between human papillomavirus and cervical cancer, as reported by the authors of the three review articles. CONCLUSION: This research demonstrated that both a technically derived and a computer derived collection, categorization, and summarization of original articles and abstracts could provide a reliable, valid, and reproducible source of ideas duplicating, to a major degree, the ideas presented by subject specialists in review articles. As such, these tools may be useful to experts preparing literature reviews by eliminating many of the clerical-mechanical features associated with present-day scientific text processing.

Bibliometrics↗

Redundant use of luminance and flashing with shape and color as highlighting codes in symbolic displays.

Three visual search experiments evaluated the benefits and distracting effects of using luminance and flashing to highlight subclasses of symbols coded by shape and color. Each of three general shape/color classes (circular/blue, diamond/red, square/yellow) was divided into three subclasses by presenting the upper half, lower half, or entire symbol. Increasing the luminance of a subclass by a factor of two did not result in a significant improvement in search performance. Flashing a subclass at a rate of 3 Hz resulted in a significantly shorter mean search time (48% improvement). Increasing the luminance of one subclass (by a factor of five) while simultaneously flashing another significantly improved search times by 31% and 43% respectively, compared with nonhighlighted search conditions. In each experiment, the search times for nonhighlighted target subclasses were not affected by the presence of brighter and flashing targets. The failure of the initial experiment to find a significant performance improvement caused by increasing symbol luminance suggested that a larger luminance increase was necessary for this code to be effective. The overall results suggest that using luminance and flashing to highlight subclasses of color- and shape-coded symbols can reduce search times for these subclasses without producing a distraction effect by way of a concomitant increase in the search times for unhighlighted symbols.

Adult↗

Disruption and maintenance of skilled visual search as a function of degree of consistency.

The present experiment was conducted to investigate the effects of varying degrees of task consistency on the performance and maintenance of skill in a semantic-category visual search task. Four groups of participants first received 6000 trials of consistent mapping (CM) training on two different categories. The participants then performed 4000 trials in which one of the previously trained categories remained 100% consistent, whereas the other previously trained category became either 100%, 67%, 50%, or 33% consistent. This second phase of the experiment allowed for the examination of disruption of the search skill as a function of degree of consistency. Subsequent to the degree of consistency manipulation, 100% consistency was restored and participants performed another 4200 CM trials. Results indicate that performance was disrupted by inconsistency and that disruption increased as consistency decreased. On the return of task consistency, performance improved rapidly to predisruption levels, though some performance disruption was evident. Theoretical and practical implications are discussed.

Adolescent↗