Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Storage And Retrieval”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

Communication and re-use of chemical information in bioscience.

The current methods of publishing chemical information in bioscience articles are analysed. Using 3 papers as use-cases, it is shown that conventional methods using human procedures, including cut-and-paste are time-consuming and introduce errors. The meaning of chemical terms and the identity of compounds is often ambiguous. valuable experimental data such as spectra and computational results are almost always omitted. We describe an Open XML architecture at proof-of-concept which addresses these concerns. Compounds are identified through explicit connection tables or links to persistent Open resources such as PubChem. It is argued that if publishers adopt these tools and protocols, then the quality and quantity of chemical information available to bioscientists will increase and the authors, publishers and readers will find the process cost-effective.

Archives↗

CGH-Profiler: data mining based on genomic aberration profiles.

BACKGROUND: CGH-Profiler is a program that supports the analysis of genomic aberrations measured by Comparative Genomic Hybridisation (CGH). Comparative genomic hybridisation (CGH) is a well-established, molecular cytogenetic method that allows the detection of chromosomal imbalances in entire genomes. This technique is widely used in routine molecular diagnostics. Typically, chromosomal imbalances are described in a complex syntax based on the International Standard for Cytogenetic Nomenclature (ISCN). This semantic description of chromosomal imbalances hinders a large-scale statistical analysis across different experiments, e.g. for finding aberration patterns associated with a particular disease type or state. RESULTS: CGH-Profiler circumvents the semantic ISCN description by importing data from different CGH system vendors and by directly transferring the data into a table format that is readily accessible for subsequent statistical analysis. CGH-profiler comes with different consistency checks, calculates various statistics and automatically assigns a median copy number ratio to each chromosomal band. Import of CGH profiles from different CGH system vendors is already supported; its extension to other systems can be readily achieved through Perl scripts.CGH profiler can also be used to analyse comparative expressed sequence hybridisation (CESH) data. CESH reveals gene expression patterns according to chromosomal locations in a similar manner as CGH detects chromosomal imbalances. CONCLUSION: CGH-Profiler is a useful tool for processing of CGH and CESH data.

Chromosome Aberrations↗

An analysis of the Sargasso Sea resource and the consequences for database composition.

BACKGROUND: The environmental sequencing of the Sargasso Sea has introduced a huge new resource of genomic information. Unlike the protein sequences held in the current searchable databases, the Sargasso Sea sequences originate from a single marine environment and have been sequenced from species that are not easily obtainable by laboratory cultivation. The resource also contains very many fragments of whole protein sequences, a side effect of the shotgun sequencing method.These sequences form a significant addendum to the current searchable databases but also present us with some intrinsic difficulties. While it is important to know whether it is possible to assign function to these sequences with the current methods and whether they will increase our capacity to explore sequence space, it is also interesting to know how current bioinformatics techniques will deal with the new sequences in the resource. RESULTS: The Sargasso Sea sequences seem to introduce a bias that decreases the potential of current methods to propose structure and function for new proteins. In particular the high proportion of sequence fragments in the resource seems to result in poor quality multiple alignments. CONCLUSION: These observations suggest that the new sequences should be used with care, especially if the information is to be used in large scale analyses. On a positive note, the results may just spark improvements in computational and experimental methods to take into account the fragments generated by environmental sequencing techniques.

Amino Acid Sequence↗

MICA: desktop software for comprehensive searching of DNA databases.

BACKGROUND: Molecular biologists work with DNA databases that often include entire genomes. A common requirement is to search a DNA database to find exact matches for a nondegenerate or partially degenerate query. The software programs available for such purposes are normally designed to run on remote servers, but an appealing alternative is to work with DNA databases stored on local computers. We describe a desktop software program termed MICA (K-Mer Indexing with Compact Arrays) that allows large DNA databases to be searched efficiently using very little memory. RESULTS: MICA rapidly indexes a DNA database. On a Macintosh G5 computer, the complete human genome could be indexed in about 5 minutes. The indexing algorithm recognizes all 15 characters of the DNA alphabet and fully captures the information in any DNA sequence, yet for a typical sequence of length L, the index occupies only about 2L bytes. The index can be searched to return a complete list of exact matches for a nondegenerate or partially degenerate query of any length. A typical search of a long DNA sequence involves reading only a small fraction of the index into memory. As a result, searches are fast even when the available RAM is limited. CONCLUSION: MICA is suitable as a search engine for desktop DNA analysis software.

Algorithms↗

Random allocation software for parallel group randomized trials.

BACKGROUND: Typically, randomization software should allow users to exert control over the different aspects of randomization including block design, provision of unique identifiers and control over the format and type of program output. While some of these characteristics have been addressed by available software, none of them have all of these capabilities integrated into one package. The main objective of the Random Allocation Software project was to enhance the user's control over different aspects of randomization in parallel group trials, including output type and format, structure and ordering of generated unique identifiers and enabling users to specify group names for more than two groups. RESULTS: The program has different settings for: simple and blocked randomizations; length, format and ordering of generated unique identifiers; type and format of program output; and saving sessions for future use. A formatted random list generated by this program can be used directly (without further formatting) by the coordinator of the research team to prepare and encode different drugs or instruments necessary for the parallel group trial. CONCLUSIONS: Random Allocation Software enables users to control different attributes of the random allocation sequence and produce qualified lists for parallel group trials.

Algorithms↗

Creating a medical English-Swedish dictionary using interactive word alignment.

BACKGROUND: This paper reports on a parallel collection of rubrics from the medical terminology systems ICD-10, ICF, MeSH, NCSP and KSH97-P and its use for semi-automatic creation of an English-Swedish dictionary of medical terminology. The methods presented are relevant for many other West European language pairs than English-Swedish. METHODS: The medical terminology systems were collected in electronic format in both English and Swedish and the rubrics were extracted in parallel language pairs. Initially, interactive word alignment was used to create training data from a sample. Then the training data were utilised in automatic word alignment in order to generate candidate term pairs. The last step was manual verification of the term pair candidates. RESULTS: A dictionary of 31,000 verified entries has been created in less than three man weeks, thus with considerably less time and effort needed compared to a manual approach, and without compromising quality. As a side effect of our work we found 40 different translation problems in the terminology systems and these results indicate the power of the method for finding inconsistencies in terminology translations. We also report on some factors that may contribute to making the process of dictionary creation with similar tools even more expedient. Finally, the contribution is discussed in relation to other ongoing efforts in constructing medical lexicons for non-English languages. CONCLUSION: In three man weeks we were able to produce a medical English-Swedish dictionary consisting of 31,000 entries and also found hidden translation errors in the utilized medical terminology systems.

Database Management Systems↗

RNdex Top 100: a quality-filtered database for nursing research.

RNdex Top 100, a value-added database providing citations and abstracts for more than 100 of the leading nursing journals, is one of the first products specifically developed for electronic searching published by SilverPlatter Information, Inc. To provide focused access to the literature, a new 9,000-term thesaurus reflecting the current terminology of the nursing profession is used. Abstracts for each record, new field descriptors, and rapid indexing are some of the features which make this database a viable alternative or supplement to CINAHL.

Abstracting and Indexing↗

ISI's Journal Citation Reports on the Web.

This column features an overview of the Institute for Scientific Information's (ISI) Journal Citation Reports (JCR) database. Basic searching techniques are presented, as well as simple ways to manipulate data contained in the file. The Journal Citation Reports database can provide information on highest impact journals, most frequently used journals, "hottest" journals, and largest journals in a field or discipline.

Academies and Institutes↗

Managing peptidases in the genomic era.

The enzymes that hydrolyse peptide bonds, called peptidases or proteases, are very important to mankind and are also very numerous. The many scientists working on these enzymes are rapidly acquiring new data, and they need good methods to store it and retrieve it. The storage and retrieval require effective systems of classification and nomenclature, and it is the design and implementation of these that we mean by 'managing' peptidases. Ten years ago Rawlings and Barrett proposed the first comprehensive system for the classification of peptidases, which included a set of simple names for the families. In the present article we describe how the system has developed since then. The peptidase classification has now been adopted for use by many other databases, and provides the structure around which the MEROPS protease database (http://merops.sanger.ac.uk) is built.

Animals↗

Which is better for presenting your data: table or graph?

This study aimed at investigating the characteristics of table and graph that people perceive and the data types which people consider the two displays are most appropriate for. Participants in this survey were 195 teachers and under-graduates from four universities in Beijing. The results showed people's different attitudes towards the two forms of display.

China↗

A new nucleotide-composition based fingerprint of SARS-CoV with visualization analysis.

It has been observed by conducting an extensive analysis of the two-dimensional cellular automata images of known SARS-CoV genome sequences that the V-shaped cross-lines only exist in some special locations, and hence can be used as a fingerprint to identify the SARS sequences. Such a discovery can be used to rapidly and reliably diagnose SARS coronavirus for both basic research in laboratories and practical application in clinics.

Algorithms↗

Scientific LogAnalyzer: a web-based tool for analyses of server log files in psychological research.

Scientific LogAnalyzer is a platform-independent interactive Web service for the analysis of log files. Scientific LogAnalyzer offers several features not available in other log file analysis tools--for example, organizational criteria and computational algorithms suited to aid behavioral and social scientists. Scientific LogAnalyzer is highly flexible on the input side (unlimited types of log file formats), while strictly keeping a scientific output format. Features include (1) free definition of log file format, (2) searching and marking dependent on any combination of strings (necessary for identifying conditions in experiment data), (3) computation of response times, (4) detection of multiple sessions, (5) speedy analysis of large log files, (6) output in HTML and/or tab-delimited form, suitable for import into statistics software, and (7) a module for analyzing and visualizing drop-out. Several methodological features specifically needed in the analysis of data collected in Internet-based experiments have been implemented in the Web-based tool and are described in this article. A regression analysis with data from 44 log file analyses shows that the size of the log file and the domain name lookup are the two main factors determining the duration of an analysis. It is less than a minute for a standard experimental study with a 2 x 2 design, a dozen Web pages, and 48 participants (ca. 800 lines, including data from drop-outs). The current version of Scientific LogAnalyzer is freely available for small log files. Its Web address is http://genpsylab-logcrunsh.unizh.ch/.

Data Interpretation, Statistical↗

Net results.

Explore the source record for details and available documents.

Computer Communication Networks↗

An EKG monitor network.

A networkable kardiomonitor CM-3 is described as well as the associated central monitoring device CEMON. CM-3 allows archiving as well as efficient control of all relevant measurements, alarms and trend data in the last 2 years of use of the equipment. This data is easily reviewed, printed or saved on removable media to be included in the hospital patient documentation. The network is based on standard Ethernet bus architecture and PC Ethernet adapters. This high speed medium allows efficient real time control and immediate reaction to each alarm situation. Easy integration with other parts of the hospital information system is possible. In addition, critical monitor files can be efficiently backed up and possibilities are open for hierarchical storage with high security options.

Computer Communication Networks↗

Self-documenting structured reports using open information standards.

Structured reporting systems use standardized data elements and predetermined data-entry formats to record observations. This article describes a system for structured data entry and reporting that generates reports encoded in the Standard Generalized Markup Language (SGML), an open, internationally accepted standard for document interchange. The structured report is self-documenting: it includes a definition of its allowable data field and values encoded as a report-specific SGML document type definition (DTD). By linking its reporting concepts with those of external vocabularies such as the UMLS Metathesaurus, this system can create open, universally comprehensible structured reports.

Data Display↗