Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Storage And Retrieval”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40Linked to original sources

Estimates of error probability for complex Gaussian channels with generalized likelihood ratio detection.

We derive approximate expressions for the probability of error in a two-class hypothesis testing problem in which the two hypotheses are characterized by zero-mean complex Gaussian distributions. These error expressions are given in terms of the moments of the test statistic employed and we derive these moments for both the likelihood ratio test, appropriate when class densities are known, and the generalized likelihood ratio test, appropriate when class densities must be estimated from training data. These moments are functions of class distribution parameters which are generally unknown so we develop unbiased moment estimators in terms of the training data. With these, accurate estimates of probability of error can be calculated quickly for both the optimal and plug-in rules from available training data. We present a detailed example of the behavior of these estimators and demonstrate their application to common pattern recognition problems, which include quantifying the incremental value of larger training data collections, evaluating relative geometry in data fusion from multiple sensors, and selecting a good subset of available features.

Algorithms↗

Learning weighted metrics to minimize nearest-neighbor classification error.

In order to optimize the accuracy of the Nearest-Neighbor classification rule, a weighted distance is proposed, along with algorithms to automatically learn the corresponding weights. These weights may be specific for each class and feature, for each individual prototype, or for both. The learning algorithms are derived by (approximately) minimizing the Leaving-One-Out classification error of the given training set. The proposed approach is assessed through a series of experiments with UCI/STATLOG corpora, as well as with a more specific task of text classification which entails very sparse data representation and huge dimensionality. In all these experiments, the proposed approach shows a uniformly good behavior, with results comparable to or better than state-of-the-art results published with the same data so far.

Algorithms↗

High-dimensional visual analytics: interactive exploration guided by pairwise views of point distributions.

We introduce a method for organizing multivariate displays and for guiding interactive exploration through high-dimensional data. The method is based on nine characterizations of the 2D distributions of orthogonal pairwise projections on a set of points in multidimensional Euclidean space. These characterizations include such measures as density, skewness, shape, outliers, and texture. Statistical analysis of these measures leads to ways for 1) organizing 2D scatterplots of points for coherent viewing, 2) locating unusual (outlying) marginal 2D distributions of points for anomaly detection, and 3) sorting multivariate displays based on high-dimensional data, such as trees, parallel coordinates, and glyphs.

Algorithms↗

Interactive visual analysis of families of function graphs.

The analysis and exploration of multidimensional and multivariate data is still one of the most challenging areas in the field of visualization. In this paper, we describe an approach to visual analysis of an especially challenging set of problems that exhibit a complex internal data structure. We describe the interactive visual exploration and analysis of data that includes several (usually large) families of function graphs fi (x, t). We describe analysis procedures and practical aspects of the interactive visual analysis specific to this type of data (with emphasis on the function graph characteristic of the data). We adopted the well-proven approach of multiple, linked views with advanced interactive brushing to assess the data. Standard views such as histograms, scatterplots, and parallel coordinates are used to jointly visualize data. We support iterative visual analysis by providing means to create complex, composite brushes that span multiple views and that are constructed using different combination schemes. We demonstrate that engineering applications represent a challenging but very applicable area for visual analytics. As a case study, we describe the optimization of a fuel injection system in diesel engines of passenger cars.

Algorithms↗

Data recording and trend display during anaesthesia using 'MacLab'.

A single screen display of variables monitored during anaesthesia may be ergonomically superior to the 'stack' of monitors seen in many anaesthetising locations. A system based on a MacLab (Analogue Digital Instruments) analogue-to-digital convertor used in conjunction with a Macintosh computer was evaluated. The system was configured to provide trend displays of up to eight variables on a single screen. It was found to be a useful adjunct to monitoring during anaesthesia. Advantages of this system are low cost, flexibility, and the quality of the software and support provided. Limitations of this and other similar systems are discussed.

Analog-Digital Conversion↗

Mining medical data.

Explore the source record for details and available documents.

Data Interpretation, Statistical↗

A phylogenomic gene cluster resource: the Phylogenetically Inferred Groups (PhIGs) database.

BACKGROUND: We present here the PhIGs database, a phylogenomic resource for sequenced genomes. Although many methods exist for clustering gene families, very few attempt to create truly orthologous clusters sharing descent from a single ancestral gene across a range of evolutionary depths. Although these non-phylogenetic gene family clusters have been used broadly for gene annotation, errors are known to be introduced by the artifactual association of slowly evolving paralogs and lack of annotation for those more rapidly evolving. A full phylogenetic framework is necessary for accurate inference of function and for many studies that address pattern and mechanism of the evolution of the genome. The automated generation of evolutionary gene clusters, creation of gene trees, determination of orthology and paralogy relationships, and the correlation of this information with gene annotations, expression information, and genomic context is an important resource to the scientific community. DISCUSSION: The PhIGs database currently contains 23 completely sequenced genomes of fungi and metazoans, containing 409,653 genes that have been grouped into 42,645 gene clusters. Each gene cluster is built such that the gene sequence distances are consistent with the known organismal relationships and in so doing, maximizing the likelihood for the clusters to represent truly orthologous genes. The PhIGs website contains tools that allow the study of genes within their phylogenetic framework through keyword searches on annotations, such as GO and InterPro assignments, and sequence similarity searches by BLAST and HMM. In addition to displaying the evolutionary relationships of the genes in each cluster, the website also allows users to view the relative physical positions of homologous genes in specified sets of genomes. SUMMARY: Accurate analyses of genes and genomes can only be done within their full phylogenetic context. The PhIGs database and corresponding website http://phigs.org address this problem for the scientific community. Our goal is to expand the content as more genomes are sequenced and use this framework to incorporate more analyses.

Base Sequence↗

Automatic extraction of candidate nomenclature terms using the doublet method.

BACKGROUND: New terminology continuously enters the biomedical literature. How can curators identify new terms that can be added to existing nomenclatures? The most direct method, and one that has served well, involves reading the current literature. The scholarly curator adds new terms as they are encountered. Present-day scholars are severely challenged by the enormous volume of biomedical literature. Curators of medical nomenclatures need computational assistance if they hope to keep their terminologies current. The purpose of this paper is to describe a method of rapidly extracting new, candidate terms from huge volumes of biomedical text. The resulting lists of terms can be quickly reviewed by curators and added to nomenclatures, if appropriate. The candidate term extractor uses a variation of the previously described doublet coding method. The algorithm, which operates on virtually any nomenclature, derives from the observation that most terms within a knowledge domain are composed entirely of word combinations found in other terms from the same knowledge domain. Terms can be expressed as sequences of overlapping word doublets that have more specific meaning than the individual words that compose the term. The algorithm parses through text, finding contiguous sequences of word doublets that are known to occur somewhere in the reference nomenclature. When a sequence of matching word doublets is encountered, it is compared with whole terms already included in the nomenclature. If the doublet sequence is not already in the nomenclature, it is extracted as a candidate new term. Candidate new terms can be reviewed by a curator to determine if they should be added to the nomenclature. An implementation of the algorithm is demonstrated, using a corpus of published abstracts obtained through the National Library of Medicine's PubMed query service and using "The developmental lineage classification and taxonomy of neoplasms" as a reference nomenclature. RESULTS: A 31+ Megabyte corpus of pathology journal abstracts was parsed using the doublet extraction method. This corpus consisted of 4,289 records, each containing an abstract title. The total number of words included in the abstract titles was 50,547. New candidate terms for the nomenclature were automatically extracted from the titles of abstracts in the corpus. Total execution time on a desktop computer with CPU speed of 2.79 GHz was 2 seconds. The resulting output consisted of 313 new candidate terms, each consisting of concatenated doublets found in the reference nomenclature. Human review of the 313 candidate terms yielded a list of 285 terms approved by a curator. A final automatic extraction of duplicate terms yielded a final list of 222 new terms (71% of the original 313 extracted candidate terms) that could be added to the reference nomenclature. CONCLUSION: The doublet method for automatically extracting candidate nomenclature terms can be used to quickly find new terms from vast amounts of text. The method can be immediately adapted for virtually any text and any nomenclature. An implementation of the algorithm, in the Perl programming language, is provided with this article.

Abstracting and Indexing↗

Information literacy and library attitudes of occupational therapy students.

Information literacy, often described as a person's ability to effectively find and evaluate answers to questions using a variety of information resources, is of particular importance to health care workers. This paper presents the results of an information literacy survey presented to occupational therapy (OT) students at Thomas Jefferson University during a series of required class activities. Also described are the authors' activities with the faculty and courses at Jefferson. The survey was made available to first-, second-, third-, and fourth-year occupational therapy students along with nursing students and pharmacy students. The survey is designed to identify research habits, skills, and preferences. Results confirm some commonly held perceptions about searching skills of young adults and an interesting dichotomy in students' learning habits. The paper concludes with a discussion of recommendations to OT faculty and librarians on how to improve information literacy education. The survey can be obtained by contacting the authors.

Adult↗

Error reduction in surgical pathology.

CONTEXT: Because of its complex nature, surgical pathology practice is inherently error prone. Currently, there is pressure to reduce errors in medicine, including pathology. OBJECTIVE: To review factors that contribute to errors and to discuss error-reduction strategies. DESIGN: Literature review. RESULTS: Multiple factors contribute to errors in medicine, including variable input, complexity, inconsistency, tight coupling, human intervention, time constraints, and a hierarchical culture. Strategies that may reduce errors include reducing reliance on memory, improving information access, error-proofing processes, decreasing reliance on vigilance, standardizing tasks and language, reducing the number of handoffs, simplifying processes, adjusting work schedules and environment, providing adequate training, and placing the correct people in the correct jobs. CONCLUSIONS: Surgical pathology is a complex system with ample opportunity for error. Significant error reduction is unlikely to occur without a sustained comprehensive program of quality control and quality assurance. Incremental adoption of information technology and automation along with improved training in patient safety and quality management can help reduce errors.

Diagnostic Errors↗

Creating user-friendly databases with Microsoft Access.

Data entry can be tedious and is fraught with potential for errors that affect study findings. Researchers can minimise entry errors and streamline data entry by using some of the popular software packages on the market. Joanne Kraenzle Schneider and colleagues describe one way to create a user-friendly database that minimises entry errors by using Microsoft (MS) Access.

Computer User Training↗