Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Datasets as Topic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Practical experience with databases for congenital heart disease: a registry versus an academic database.

Increasingly, pooled data from multiple institutions are the source of published clinical results. A computerized database program is essential to compile and analyze clinical experience. The scope of data collection defines a database. Two types of databases, the registry and academic, are compared. In a registry database, some of the data are collected on all patients. The resources dedicated to data collection and entry are the practical limit to the extent of information in the database. The agreement on nomenclature for surgical diagnosis and procedure codes of congenital heart disease has paved the way for the development of a multi-institutional registry database. The registry database could provide a standard of care reference for early results after congenital heart surgery. The practical difficulty of data collection is obviated by limiting information to a basic minimum dataset. The academic database, in which all of the data are collected for a defined subset of patients, is designed to investigate a specific population of patients to generate new knowledge. It contains sufficient data to allow sophisticated statistical analysis to clarify the determinants of good and poor outcome, including early, mid- and long-term follow-up information. Multi-institutional pooling of detailed information derived from academic databases will be of increasing importance in generating new knowledge to foster improved therapy for patients with congenital heart disease.

Cardiac Surgical Procedures↗

Protein interactions: two methods for assessment of the reliability of high throughput observations.

High throughput methods for detecting protein interactions require assessment of their accuracy. We present two forms of computational assessment. The first method is the expression profile reliability (EPR) index. The EPR index estimates the biologically relevant fraction of protein interactions detected in a high throughput screen. It does so by comparing the RNA expression profiles for the proteins whose interactions are found in the screen with expression profiles for known interacting and non-interacting pairs of proteins. The second form of assessment is the paralogous verification method (PVM). This method judges an interaction likely if the putatively interacting pair has paralogs that also interact. In contrast to the EPR index, which evaluates datasets of interactions, PVM scores individual interactions. On a test set, PVM identifies correctly 40% of true interactions with a false positive rate of approximately 1%. EPR and PVM were applied to the Database of Interacting Proteins (DIP), a large and diverse collection of protein-protein interactions that contains over 8000 Saccharomyces cerevisiae pairwise protein interactions. Using these two methods, we estimate that approximately 50% of them are reliable, and with the aid of PVM we identify confidently 3003 of them. Web servers for both the PVM and EPR methods are available on the DIP website (dip.doe-mbi.ucla.edu/Services.cgi).

Algorithms↗

Exploring the altered daily geographies and lifeworlds of women living with fibromyalgia syndrome: a mixed-method approach.

In this paper I employ data triangulation in order to investigate the complex nature of the altered lifeworlds and daily geographies of women living with fibromyalgia syndrome (FMS). More specifically, I use the findings of in-depth interviews and a standardized test (the Sickness Impact Profile [SIP]) in a mixed-method approach to understanding how women's lives change after the onset of FMS and how their changing bodies and locations in society and space shape such altered lifeworlds. These data were collected from 55 women living with FMS in Ontario, Canada. The experiential evidence shared during the interviews is used to qualify or explain certain phenomena observed within the SIP dataset. I focus on four specific experiences in the women's lives; these are the: (1) onset of mental haziness and fatigue; (2) development of disrupted sleep/sleep disorders; (3) removal from paid labour; and (4) withdrawal from social and recreational activities. It is found that changes in the women's bodies precipitated some of the most significant life changes experienced, including altered identities and diminished incomes, and that altered bodily realities facilitated or denied access to socio-spatial life. At the same time, the women's changing locations in society and space also played a role in bringing about such changes.

Cost of Illness↗

A body image scale for use with cancer patients.

Body image is an important endpoint in quality of life evaluation since cancer treatment may result in major changes to patients' appearance from disfiguring surgery, late effects of radiotherapy or adverse effects of systemic treatment. A need was identified to develop a short body image scale (BIS) for use in clinical trials. A 10-item scale was constructed in collaboration with the European Organization for Research and Treatment of Cancer (EORTC) Quality of Life Study Group and tested in a heterogeneous sample of 276 British cancer patients. Following revisions, the scale underwent psychometric testing in 682 patients with breast cancer, using datasets from seven UK treatment trials/clinical studies. The scale showed high reliability (Cronbach's alpha 0.93) and good clinical validity based on response prevalence, discriminant validity (P<0.0001, Mann-Whitney test), sensitivity to change (P<0.001, Wilcoxon signed ranks test) and consistency of scores from different breast cancer treatment centres. Factor analysis resulted in a single factor solution in three out of four analyses, accounting for >50% variance. These results support the clinical validity of the BIS as a brief questionnaire for assessing body image changes in patients with cancer, suitable for use in clinical trials.

Age Factors↗

Correlation between gene expression profiles and protein-protein interactions within and across genomes.

MOTIVATION: Function annotation of an unclassified protein on the basis of its interaction partners is well documented in the literature. Reliable predictions of interactions from other data sources such as gene expression measurements would provide a useful route to function annotation. We investigate the global relationship of protein-protein interactions with gene expression. This relationship is studied in four evolutionarily diverse species, for which substantial information regarding their interactions and expression is available: human, mouse, yeast and Escherichia coli. RESULTS: In E.coli the expression of interacting pairs is highly correlated in comparison to random pairs, while in the other three species, the correlation of expression of interacting pairs is only slightly stronger than that of random pairs. To strengthen the correlation, we developed a protocol to integrate ortholog information into the interaction and expression datasets. In all four genomes, the likelihood of predicting protein interactions from highly correlated expression data is increased using our protocol. In yeast, for example, the likelihood of predicting a true interaction, when the correlation is > 0.9, increases from 1.4 to 9.4. The improvement demonstrates that protein interactions are reflected in gene expression and the correlation between the two is strengthened by evolution information. The results establish that co-expression of interacting protein pairs is more conserved than that of random ones.

Animals↗

Sample size and power estimation for studies with health related quality of life outcomes: a comparison of four methods using the SF-36.

We describe and compare four different methods for estimating sample size and power, when the primary outcome of the study is a Health Related Quality of Life (HRQoL) measure. These methods are: 1. assuming a Normal distribution and comparing two means; 2. using a non-parametric method; 3. Whitehead's method based on the proportional odds model; 4. the bootstrap. We illustrate the various methods, using data from the SF-36. For simplicity this paper deals with studies designed to compare the effectiveness (or superiority) of a new treatment compared to a standard treatment at a single point in time. The results show that if the HRQoL outcome has a limited number of discrete values (< 7) and/or the expected proportion of cases at the boundaries is high (scoring 0 or 100), then we would recommend using Whitehead's method (Method 3). Alternatively, if the HRQoL outcome has a large number of distinct values and the proportion at the boundaries is low, then we would recommend using Method 1. If a pilot or historical dataset is readily available (to estimate the shape of the distribution) then bootstrap simulation (Method 4) based on this data will provide a more accurate and reliable sample size estimate than conventional methods (Methods 1, 2, or 3). In the absence of a reliable pilot set, bootstrapping is not appropriate and conventional methods of sample size estimation or simulation will need to be used. Fortunately, with the increasing use of HRQoL outcomes in research, historical datasets are becoming more readily available. Strictly speaking, our results and conclusions only apply to the SF-36 outcome measure. Further empirical work is required to see whether these results hold true for other HRQoL outcomes. However, the SF-36 has many features in common with other HRQoL outcomes: multi-dimensional, ordinal or discrete response categories with upper and lower bounds, and skewed distributions, so therefore, we believe these results and conclusions using the SF-36 will be appropriate for other HRQoL measures.

Algorithms↗

A grid-based image archival and analysis system.

Here the authors present a Grid-aware middleware system, called GridPACS, that enables management and analysis of images in a massive scale, leveraging distributed software components coupled with interconnected computation and storage platforms. The need for this infrastructure is driven by the increasing biomedical role played by complex datasets obtained through a variety of imaging modalities. The GridPACS architecture is designed to support a wide range of biomedical applications encountered in basic and clinical research, which make use of large collections of images. Imaging data yield a wealth of metabolic and anatomic information from macroscopic (e.g., radiology) to microscopic (e.g., digitized slides) scale. Whereas this information can significantly improve understanding of disease pathophysiology as well as the noninvasive diagnosis of disease in patients, the need to process, analyze, and store large amounts of image data presents a great challenge.

Computer Communication Networks↗

[Isolated fractures of the orbital floor].

BACKGROUND: The goal of this retrospective study was quantitative calculation of area and volume of isolated orbital floor fractures from computed tomography (CT) and correlation of these data with post-traumatic ophthalmologic findings. PATIENTS AND METHODS: A total of 76 patients with isolated orbital floor fractures were evaluated radiologically and clinically. CT scanning was performed in coronal sections (1.5-mm to 3.0-mm slice thickness) with contiguous table feed. Orbital floor and fracture area as well as volume of displaced tissue were measured and calculated from the CT dataset. The relation of quantitative CT data to ophthalmologic findings (motility, diplopia, and globe position) was assessed statistically. RESULTS: Calculation of the CT dataset revealed a mean orbital floor area of 6.33+/-1.05 cm(2), a mean fracture area of 2.60+/-1.14 cm(2), and a mean volume of displaced tissue of 1.16+/-0.80 cm(3). Volume of displaced tissue correlated significantly with ophthalmologic findings (p< or =0.01). Fracture area correlated significantly with globe position (p< or =0.01) and was less associated with diplopia and motility disturbances (p<0.10). CONCLUSION: Efficient evaluation of two-dimensional CT data enables quantitative assessment of orbital floor fractures. Position and function of the globe are mainly affected by the volume of displaced periorbital tissue.

Adult↗

Improved FlexX docking using FlexS-determined base fragment placement.

We report on a novel hybrid FlexX/FlexS docking approach, whereby the base fragment of the test ligand is chosen by FlexS superposition onto a cocrystallized template ligand and then fed into FlexX for the incremental construction of the final solution. The new approach is tested on the diverse 200 protein-ligand complex dataset that has been previously described for FlexX validation. In total, 62.9% of the complexes can be reproduced at rank 1 by our approach, which compares favorably with 46.9% when using FlexX alone. In addition, we report "cross-docking" experiments in which several receptor structures of complexes with identical proteins have been used for docking all cocrystallized ligands of these complexes. The results show that, in almost all cases, the hybrid approach can acceptably dock a ligand into a foreign receptor structure using a different ligand template, can give solutions where FlexX alone fails, and tends to give solutions that are more accurately positioned.

Algorithms↗

A new method for computing the multipoint posterior probability of linkage.

The posterior probability of linkage (PPL) is a Bayesian statistic which directly measures the probability of linkage between a trait locus and a marker (in the 2-point case) or a genomic region (in the multipoint case). It has several benefits, including ease of interpretation, the ability to incorporate prior genomic information, and a mathematically rigorous and robust procedure for accumulating linkage information across multiple heterogeneous datasets. To date, the majority of work on the PPL has focused on the development of the 2-point statistic, with only preliminary attempts at the development of an equivalent multipoint version. In this paper we present a new way of computing of the multipoint PPL. This new version imputes to each genomic point an estimate of the 2-point PPL we would have obtained from a fully informative marker giving similar evidence for linkage. This version, which we call the imputed PPL, is shown to be superior to previously developed versions.

Bayes Theorem↗

Comparison of intron-containing and intron-lacking human genes elucidates putative exonic splicing enhancers.

Of the rules used by the splicing machinery to precisely determine intron-exon boundaries only a fraction is known. Recent evidence suggests that specific short sequences within exons help in defining these boundaries. Such sequences are known as exonic splicing enhancers (ESE). A possible bioinformatical approach to studying ESE sequences is to compare genes that harbor introns with genes that do not. For this purpose two non-redundant samples of 719 intron-containing and 63 intron-lacking human genes were created. We performed a statistical analysis on these datasets of intron-containing and intron-lacking human coding sequences and found a statistically significant difference (P = 0.01) between these samples in terms of 5-6mer oligonucleotide distributions. The difference is not created by a few strong signals present in the majority of exons, but rather by the accumulation of multiple weak signals through small variations in codon frequencies, codon biases and context-dependent codon biases between the samples. A list of putative novel human splicing regulation sequences has been elucidated by our analysis.

Alternative Splicing↗

Safety, permanency, and in-home services: applying administrative data.

This article describes the construction and use of safety and permanency indicators, two aspects of a full set of indicators that also includes child well-being and family functioning. The indicators were constructed from Philadelphia's Family and Child Tracking System and were used to examine the city's Services to Children in their Own Home (SCOH) program. Cohort datasets were constructed through the use of extract files, and two independent data file construction algorithms were employed to calibrate the accuracy of the data construction process. The primary unit of analysis was the "family" spell in SCOH services. Contextual variables included family structure, race, and service intensity. The indicators associated with SCOH spells included reports of maltreatment after service, founded maltreatment after service, and out-of-home placement after service. Event history techniques were used to conduct the data analysis. Baseline indicator data for Philadelphia are presented, and future uses for such data are discussed.

Child↗

NMR spectral quantitation by principal-component analysis. II. Determination of frequency and phase shifts.

This paper extends the use of principal-component analysis in spectral quantification to the estimation of frequency and phase shifts in a single resonant peak across a series of spectra. The estimated parameters can be used to correct the spectra accordingly, resulting in more accurate peak-area estimation. Further, the removal of the variations in phase and frequency cause by instrumental and experimental fluctuations makes it possible to determine more accurately the remaining variations, which bear biological significance. The procedure is demonstrated on simulated data, a 3D chemical-shift-imaging dataset acquired from a cylinder of inorganic phosphate (Pi), and a set of 736 31P NMR in vivo spectra taken from a kinetic study of rate muscle energetics. In all cases, the procedure rapidly and automatically identifies the frequency and phase shifts present in the individual spectra. In the kinetic study, the procedure is used twice, first to adjust the phase and frequency of a reference peak (phosphocreatine) and then to determine the individual frequencies of the Pi peak in each of the spectra which further can be used for estimation of pH changes during the experiment.

Computer Simulation↗

Genes, age, and alcoholism: analysis of GAW14 data.

A genetic analysis of age of onset of alcoholism was performed on the Collaborative Study on the Genetics of Alcoholism data released for Genetic Analysis Workshop 14. Our study illustrates an application of the log-normal age of onset model in our software Genetic Epidemiology Models (GEMs). The phenotype ALDX1 of alcoholism was studied. The analysis strategy was to first find the markers of the Affymetrix SNP dataset with significant association with age of onset, and then to perform linkage analysis on them. ALDX1 revealed strong evidence of linkage for marker tsc0041591 on chromosome 2 and suggestive linkage for marker tsc0894042 on chromosome 3. The largest separation in mean ages of onset of ALDX1 was 19.76 and 24.41 between male smokers who are carriers of the risk allele of tsc0041591 and the non-carriers, respectively. Hence, male smokers who are carriers of marker tsc0041591 on chromosome 2 have an average onset of ALDX1 almost 5 years earlier than non-carriers.

Age of Onset↗

B-SPID: an object-relational database architecture to store, retrieve, and manipulate neuroimaging data.

We propose a hardware and software architecture to respond to crucial problems in the neuroimaging field: storage, retrieval, and processing of large datasets. The B-SPID project, here discussed, concerns the processing of neuroimages and attached components stored in an object-relational multimedia database management system (DBMS). Advanced bioinformation concepts are exploited in this project such as large scale data storage, high level graphical user interfaces and 3D graphical processing and display of data. Our database implementation is based on standard programming components, runs on several UNIX platforms and is written to be evolutive. Queries on this database are designed to obtain and display from neuroimaging data several types of results (pictures, text, or 3D graphical shapes) on heterogeneous systems.

Brain Mapping↗

Molecular phylogeny and evolutionary history of the tit-tyrants (Aves: Tyrannidae).

Tit-tyrants of the genus Anairetes presently consist of six species; five inhabit various regions along the Andean cordillera of South America and one is endemic to the Juan Fernandez Islands off the coast of Chile. Data from mtDNA ND2 and Cyt b sequences were used to construct a phylogeny for all Anairetes species as well as Uromyias agilis, a closely related genus, and Stigmatura as an outgroup, to determine their relationships and history of radiation in South America. Results strongly supported the following paired relationships: A. nigrocristatus-A. reguloides, A. flavirostris-A. alpinus, and A. parulus-A. fernandezianus. This dataset, however, could not resolve basal nodes; therefore relationships among these pairs remains obscure. Moreover the genus Uromyias, controversially separated on morphological criteria from Anairetes, fell within the Anairetes clade, although its exact position could not be ascertained with confidence. The molecular data indicate that this group probably radiated within the past 2 million years, concomitant with highly accentuated cycles of global climatic change. Certain high altitude areas within the Andes may have been stable during global climatic changes and may have served as refugia during the Plio-Pleistocene.

Animals↗

Nomenclature-based data retrieval without prior annotation: facilitating biomedical data integration with fast doublet matching.

Assigning nomenclature codes to biomedical data is an arduous, expensive and error-prone task. Data records are coded to to provide a common representation of contained concepts, allowing facile retrieval of records via a standard terminology. In the medical field, cancer registrars, nurses, pathologists, and private clinicians all understand the importance of annotating medical records with vocabularies that codify the names of diseases, procedures, billing categories, etc. Molecular biologists need codified medical records so that they can discover or validate relationships between experimental data and clinical data. This paper introduces a new approach to retrieving data records without prior coding. The approach achieves the same result as a search over pre-coded records. It retrieves all records that contain any terms that are synonymous with a user's query-term. A recently described fast algorithm (the doublet method) permits quick iterative searches over every synonym for any term from any nomenclature occurring in a dataset of any size. As a demonstration, a 105+ Megabyte corpus of Pubmed abstracts was searched for medical terms. Query terms were matched against either of two vocabularies and expanded as an array of equivalent search items. A single search term may have over one hundred nomenclature synonyms, all of which were searched against the full database. Iterative searches of a list of concept-equivalent terms involves many more operations than a single search over pre-annotated concept codes. Nonetheless, the doublet method achieved fast query response times (0.05 seconds using Snomed and 5 seconds using the Developmental Lineage Classification of Neoplasms, on a computer with a 2.89 GHz processor). Pre-annotated datasets lose their value when the chosen vocabulary is replaced by a different vocabulary or by a different version of the same vocabulary. The doublet method can employ any version of any vocabulary with no pre-annotation. In many instances, the enormous effort and expense associated with data annotation can be eliminated by on-the-fly doublet matching. The algorithm for nomenclature-based database searches using the doublet method is described. Perl scripts for implementing the algorithm and testing execution speed are provided as open source documents available from the Association for Pathology Informatics (www.pathologyinformatics.org/informatics_r.htm).

Abstracting and Indexing↗

Comparison of chemotherapy and bone marrow transplants using two independent clinical databases.

Comparing the outcome of chemotherapy and bone marrow transplants in the absence of a randomized trial is difficult but necessary for diseases where small numbers of patients make such trials difficult if not impossible. To address this issue for adults with acute lymphoblastic leukemia in first remission, we created an empirical database using two separate datasets, one from the International Bone Marrow Transplant Registry and the other from two multicenter chemotherapy studies. Prior to combining the datasets, a study protocol was developed to define inclusion criteria, outcomes to be compared and statistical methods. The main problems of a non-randomized comparison are biases potentially introduced by differences in baseline composition of the two cohorts and differences in time-to-treatment. The source of the latter bias is different distributions of waiting times between achieving complete remission and receiving post-remission therapy. Several techniques to control these biases were evaluated; each gave qualitatively similar results. These methods can easily be applied to other clinical situations where randomized trials are not available.

Adolescent↗