Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Datasets as Topic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

[Isolated fractures of the orbital floor].

BACKGROUND: The goal of this retrospective study was quantitative calculation of area and volume of isolated orbital floor fractures from computed tomography (CT) and correlation of these data with post-traumatic ophthalmologic findings. PATIENTS AND METHODS: A total of 76 patients with isolated orbital floor fractures were evaluated radiologically and clinically. CT scanning was performed in coronal sections (1.5-mm to 3.0-mm slice thickness) with contiguous table feed. Orbital floor and fracture area as well as volume of displaced tissue were measured and calculated from the CT dataset. The relation of quantitative CT data to ophthalmologic findings (motility, diplopia, and globe position) was assessed statistically. RESULTS: Calculation of the CT dataset revealed a mean orbital floor area of 6.33+/-1.05 cm(2), a mean fracture area of 2.60+/-1.14 cm(2), and a mean volume of displaced tissue of 1.16+/-0.80 cm(3). Volume of displaced tissue correlated significantly with ophthalmologic findings (p< or =0.01). Fracture area correlated significantly with globe position (p< or =0.01) and was less associated with diplopia and motility disturbances (p<0.10). CONCLUSION: Efficient evaluation of two-dimensional CT data enables quantitative assessment of orbital floor fractures. Position and function of the globe are mainly affected by the volume of displaced periorbital tissue.

Adult↗

Improved FlexX docking using FlexS-determined base fragment placement.

We report on a novel hybrid FlexX/FlexS docking approach, whereby the base fragment of the test ligand is chosen by FlexS superposition onto a cocrystallized template ligand and then fed into FlexX for the incremental construction of the final solution. The new approach is tested on the diverse 200 protein-ligand complex dataset that has been previously described for FlexX validation. In total, 62.9% of the complexes can be reproduced at rank 1 by our approach, which compares favorably with 46.9% when using FlexX alone. In addition, we report "cross-docking" experiments in which several receptor structures of complexes with identical proteins have been used for docking all cocrystallized ligands of these complexes. The results show that, in almost all cases, the hybrid approach can acceptably dock a ligand into a foreign receptor structure using a different ligand template, can give solutions where FlexX alone fails, and tends to give solutions that are more accurately positioned.

Algorithms↗

A new method for computing the multipoint posterior probability of linkage.

The posterior probability of linkage (PPL) is a Bayesian statistic which directly measures the probability of linkage between a trait locus and a marker (in the 2-point case) or a genomic region (in the multipoint case). It has several benefits, including ease of interpretation, the ability to incorporate prior genomic information, and a mathematically rigorous and robust procedure for accumulating linkage information across multiple heterogeneous datasets. To date, the majority of work on the PPL has focused on the development of the 2-point statistic, with only preliminary attempts at the development of an equivalent multipoint version. In this paper we present a new way of computing of the multipoint PPL. This new version imputes to each genomic point an estimate of the 2-point PPL we would have obtained from a fully informative marker giving similar evidence for linkage. This version, which we call the imputed PPL, is shown to be superior to previously developed versions.

Bayes Theorem↗

Comparison of intron-containing and intron-lacking human genes elucidates putative exonic splicing enhancers.

Of the rules used by the splicing machinery to precisely determine intron-exon boundaries only a fraction is known. Recent evidence suggests that specific short sequences within exons help in defining these boundaries. Such sequences are known as exonic splicing enhancers (ESE). A possible bioinformatical approach to studying ESE sequences is to compare genes that harbor introns with genes that do not. For this purpose two non-redundant samples of 719 intron-containing and 63 intron-lacking human genes were created. We performed a statistical analysis on these datasets of intron-containing and intron-lacking human coding sequences and found a statistically significant difference (P = 0.01) between these samples in terms of 5-6mer oligonucleotide distributions. The difference is not created by a few strong signals present in the majority of exons, but rather by the accumulation of multiple weak signals through small variations in codon frequencies, codon biases and context-dependent codon biases between the samples. A list of putative novel human splicing regulation sequences has been elucidated by our analysis.

Alternative Splicing↗

Safety, permanency, and in-home services: applying administrative data.

This article describes the construction and use of safety and permanency indicators, two aspects of a full set of indicators that also includes child well-being and family functioning. The indicators were constructed from Philadelphia's Family and Child Tracking System and were used to examine the city's Services to Children in their Own Home (SCOH) program. Cohort datasets were constructed through the use of extract files, and two independent data file construction algorithms were employed to calibrate the accuracy of the data construction process. The primary unit of analysis was the "family" spell in SCOH services. Contextual variables included family structure, race, and service intensity. The indicators associated with SCOH spells included reports of maltreatment after service, founded maltreatment after service, and out-of-home placement after service. Event history techniques were used to conduct the data analysis. Baseline indicator data for Philadelphia are presented, and future uses for such data are discussed.

Child↗

NMR spectral quantitation by principal-component analysis. II. Determination of frequency and phase shifts.

This paper extends the use of principal-component analysis in spectral quantification to the estimation of frequency and phase shifts in a single resonant peak across a series of spectra. The estimated parameters can be used to correct the spectra accordingly, resulting in more accurate peak-area estimation. Further, the removal of the variations in phase and frequency cause by instrumental and experimental fluctuations makes it possible to determine more accurately the remaining variations, which bear biological significance. The procedure is demonstrated on simulated data, a 3D chemical-shift-imaging dataset acquired from a cylinder of inorganic phosphate (Pi), and a set of 736 31P NMR in vivo spectra taken from a kinetic study of rate muscle energetics. In all cases, the procedure rapidly and automatically identifies the frequency and phase shifts present in the individual spectra. In the kinetic study, the procedure is used twice, first to adjust the phase and frequency of a reference peak (phosphocreatine) and then to determine the individual frequencies of the Pi peak in each of the spectra which further can be used for estimation of pH changes during the experiment.

Computer Simulation↗

B-SPID: an object-relational database architecture to store, retrieve, and manipulate neuroimaging data.

We propose a hardware and software architecture to respond to crucial problems in the neuroimaging field: storage, retrieval, and processing of large datasets. The B-SPID project, here discussed, concerns the processing of neuroimages and attached components stored in an object-relational multimedia database management system (DBMS). Advanced bioinformation concepts are exploited in this project such as large scale data storage, high level graphical user interfaces and 3D graphical processing and display of data. Our database implementation is based on standard programming components, runs on several UNIX platforms and is written to be evolutive. Queries on this database are designed to obtain and display from neuroimaging data several types of results (pictures, text, or 3D graphical shapes) on heterogeneous systems.

Brain Mapping↗

Molecular phylogeny and evolutionary history of the tit-tyrants (Aves: Tyrannidae).

Tit-tyrants of the genus Anairetes presently consist of six species; five inhabit various regions along the Andean cordillera of South America and one is endemic to the Juan Fernandez Islands off the coast of Chile. Data from mtDNA ND2 and Cyt b sequences were used to construct a phylogeny for all Anairetes species as well as Uromyias agilis, a closely related genus, and Stigmatura as an outgroup, to determine their relationships and history of radiation in South America. Results strongly supported the following paired relationships: A. nigrocristatus-A. reguloides, A. flavirostris-A. alpinus, and A. parulus-A. fernandezianus. This dataset, however, could not resolve basal nodes; therefore relationships among these pairs remains obscure. Moreover the genus Uromyias, controversially separated on morphological criteria from Anairetes, fell within the Anairetes clade, although its exact position could not be ascertained with confidence. The molecular data indicate that this group probably radiated within the past 2 million years, concomitant with highly accentuated cycles of global climatic change. Certain high altitude areas within the Andes may have been stable during global climatic changes and may have served as refugia during the Plio-Pleistocene.

Animals↗

Nomenclature-based data retrieval without prior annotation: facilitating biomedical data integration with fast doublet matching.

Assigning nomenclature codes to biomedical data is an arduous, expensive and error-prone task. Data records are coded to to provide a common representation of contained concepts, allowing facile retrieval of records via a standard terminology. In the medical field, cancer registrars, nurses, pathologists, and private clinicians all understand the importance of annotating medical records with vocabularies that codify the names of diseases, procedures, billing categories, etc. Molecular biologists need codified medical records so that they can discover or validate relationships between experimental data and clinical data. This paper introduces a new approach to retrieving data records without prior coding. The approach achieves the same result as a search over pre-coded records. It retrieves all records that contain any terms that are synonymous with a user's query-term. A recently described fast algorithm (the doublet method) permits quick iterative searches over every synonym for any term from any nomenclature occurring in a dataset of any size. As a demonstration, a 105+ Megabyte corpus of Pubmed abstracts was searched for medical terms. Query terms were matched against either of two vocabularies and expanded as an array of equivalent search items. A single search term may have over one hundred nomenclature synonyms, all of which were searched against the full database. Iterative searches of a list of concept-equivalent terms involves many more operations than a single search over pre-annotated concept codes. Nonetheless, the doublet method achieved fast query response times (0.05 seconds using Snomed and 5 seconds using the Developmental Lineage Classification of Neoplasms, on a computer with a 2.89 GHz processor). Pre-annotated datasets lose their value when the chosen vocabulary is replaced by a different vocabulary or by a different version of the same vocabulary. The doublet method can employ any version of any vocabulary with no pre-annotation. In many instances, the enormous effort and expense associated with data annotation can be eliminated by on-the-fly doublet matching. The algorithm for nomenclature-based database searches using the doublet method is described. Perl scripts for implementing the algorithm and testing execution speed are provided as open source documents available from the Association for Pathology Informatics (www.pathologyinformatics.org/informatics_r.htm).

Abstracting and Indexing↗

Comparison of chemotherapy and bone marrow transplants using two independent clinical databases.

Comparing the outcome of chemotherapy and bone marrow transplants in the absence of a randomized trial is difficult but necessary for diseases where small numbers of patients make such trials difficult if not impossible. To address this issue for adults with acute lymphoblastic leukemia in first remission, we created an empirical database using two separate datasets, one from the International Bone Marrow Transplant Registry and the other from two multicenter chemotherapy studies. Prior to combining the datasets, a study protocol was developed to define inclusion criteria, outcomes to be compared and statistical methods. The main problems of a non-randomized comparison are biases potentially introduced by differences in baseline composition of the two cohorts and differences in time-to-treatment. The source of the latter bias is different distributions of waiting times between achieving complete remission and receiving post-remission therapy. Several techniques to control these biases were evaluated; each gave qualitatively similar results. These methods can easily be applied to other clinical situations where randomized trials are not available.

Adolescent↗

Integrated statistical analysis of cDNA microarray and NIR spectroscopic data applied to a hemp dataset.

Both cDNA microarray and spectroscopic data provide indirect information about the chemical compounds present in the biological tissue under consideration. In this paper simple univariate and bivariate measures are used to investigate correlations between both types of high dimensional analyses. A large dataset of 42 hemp samples on which 3456 cDNA clones and 351 NIR wavelengths have been measured, was analyzed using graphical representations. For this purpose we propose clustered correlation and clustered discrimination images. Large, tissue-related differences are seen to dominate the cDNA-NIR correlation structure but smaller, more difficult to detect, variety-related differences can be found at specific cDNA clone/NIR wavelength combinations.

Algorithms↗

Profound effect of normalization on detection of differentially expressed genes in oligonucleotide microarray data analysis.

BACKGROUND: Oligonucleotide microarrays measure the relative transcript abundance of thousands of mRNAs in parallel. A large number of procedures for normalization and detection of differentially expressed genes have been proposed. However, the relative impact of these methods on the detection of differentially expressed genes remains to be determined. RESULTS: We have employed four different normalization methods and all possible combinations with three different statistical algorithms for detection of differentially expressed genes on a prototype dataset. The number of genes detected as differentially expressed differs by a factor of about three. Analysis of lists of genes detected as differentially expressed, and rank correlation coefficients for probability of differential expression shows that a high concordance between different methods can only be achieved by using the same normalization procedure. CONCLUSIONS: Normalization has a profound influence of detection of differentially expressed genes. This influence is higher than that of three subsequent statistical analysis procedures examined. Algorithms incorporating more array-derived information than gene-expression values alone are urgently needed.

Animals↗

PPD - Proteome Profile Database.

With the complete sequencing of multiple genomes, there have been extensions in the methods of sequence analysis from single gene/protein-based to analyzing multiple genes and proteins simultaneously. Therefore, there is a demand of user-friendly software tools that will allow mining of these enormous datasets. PPD is a WWW-based database for comparative analysis of protein lengths in completely sequenced prokaryotic and eukaryotic genomes. PPD's core objective is to create protein classification tables based on the lengths of proteins by specifying a set of organisms and parameters. The interface can also generate information on changes in proteins of specific length distributions. This feature is of importance when the user's interest is focused on some evolutionarily related organisms or on organisms with similar or related tissue specificity or life-style. PPD is available at: PPD Home.

Animals↗

The classification of subjects with joint complaints on incomplete biochemical and haematological datasets.

We performed a retrospective study on 163 subjects suffering from rheumatic fever (16), rheumatoid arthritis (36), lupus erythematosus (17), gout (21), arthrosis (50) and osteomyelitis (23). The number of variables evaluated was 39. These were all of a general biochemical and haematological nature. A feature reduction resulted in sixteen variables that matched well with those known from the literature. Linear discriminant analysis yielded poor results in classifying the six disease categories (with 18 variables 61.8%). A reduction to three disease categories improved the classification results remarkably. This, and the excellent discriminating power between patients and the reference group, shows that the selected variables are illustrative only for general clinical pictures, such as infection, and not for the desired differential diagnosis.

Arthritis, Rheumatoid↗

Dupuytren's disease risk factors.

Dupuytren's is a common problem, but little is known about its aetiology. We have undertaken a large case-control study to assess and quantify the relative contributions of diabetes and epilepsy as risk factors for Dupuytren's in the community. Cases were patients with a diagnosis of Dupuytren's disease and, for each, two controls were individually matched by age, sex, and general practice. Our dataset included 821 cases and 1,642 controls. Five hundred and eighty-eight (72%) of the cases were men. The mean age at diagnosis was 62 (range 24-97) years. Diabetes was a significant risk factor for Dupuytren's disease (OR=1.75) and there was an increased risk for medicinally treated diabetes (metformin--OR=3.56; sulphonylureas--OR=1.75) and particularly insulin controlled (OR=4.38) rather than diet-controlled diabetes. Epilepsy (OR=1.12) and anti-epileptic medications were not associated with Dupuytren's disease. Ascertainment bias in previous studies may explain the reported association with epilepsy.

Adult↗

Multi-species assessment of electrical resistance as a skin integrity marker for in vitro percutaneous absorption studies.

Assessment of percutaneous absorption in vitro provides key information when predicting dermal absorption in vivo. Confirmation of skin membrane integrity is an essential component of the in vitro method, as described in test guideline OECD 428. Historically, assessment of the membrane's permeability to tritiated water (T2O) and the generation of a permeability coefficient (Kp) were used to confirm that the skin membrane was intact prior to application of the test penetrant. Measuring electrical resistance (ER) across the membrane is a simpler, quicker, safer and more cost effective method. To investigate the robustness of the ER integrity measure, the Kp values for T2O for a range of human and animal skin membranes were compared with corresponding ER data. Overall, for human, rat, pig, mouse, rabbit and guinea pig skin, the ER data gave a good inverse association with the corresponding Kp values; the higher the Kp the lower the ER values. In addition, the distribution across a large dataset for individual skin samples was similar for Kp and ER, allowing a cut-off value for ER to be established for each skin type. Based on CTL's (Syngenta Central Toxicology Laboratory) standard static diffusion cells and databridge, we propose that intact skin should have an ER equal to or above (in kOmega): human (10), mouse (5) guinea pig (5), pig (4) rat (3), and rabbit (0.8). We conclude that measurement of ER across in vitro skin membranes provides a robust measurement of skin barrier integrity and is an appropriate alternative to Kp for T2O in order to identify intact membranes that have acceptable permeability characteristics for in vitro percutaneous absorption studies.

Administration, Topical↗

Serotonergic brainstem abnormalities in Northern Plains Indians with the sudden infant death syndrome.

The rate of the sudden infant death syndrome (SIDS) among American Indian infants in the Northern Plains is almost 6 times higher than in U.S. white infants. In a study of infant mortality among Northern Plains Indians, we tested the hypothesis that receptor binding abnormalities to the neurotransmitter serotonin (5-HT) in SIDS cases, compared with autopsied controls, occur in regions of the medulla oblongata that contain 5-HT neurons and that are critical for the regulation of cardiorespiration and central chemosensitivity during sleep, i.e. the medullary 5-HT system. Tritiated-lysergic acid diethylamide binding to 5-HT(1A-D) and 5-HT2 receptors was measured in 19 brainstem nuclei in 23 SIDS and 6 control infants using tissue receptor autoradiography. Binding in the arcuate nucleus, a part of the medullary 5-HT system along the ventral surface, in the SIDS infants (mean age-adjusted binding 7.1 +/- 0.8 fmol/mg tissue, n = 23) was significantly lower than in controls (mean age-adjusted binding 13.1 +/- 1.6 fmol/mg tissue, n = 5) (p = 0.003). Binding also demonstrated significant diagnosis x age interactions (p < 0.04) in 4 other nuclei that are components of the 5-HT system. These data suggest that medullary 5-HT dysfunction can lead to sleep-related, sudden death in affected SIDS infants, and confirm the same binding abnormalities reported by us in a larger dataset of non-American Indian SIDS and control infants. This study also links 5-HT abnormalities in the arcuate nucleus with exposure to adverse prenatal exposures, i.e. cigarette smoking (p = 0.011) and alcohol (p = 0.075), during the periconceptional period or throughout pregnancy. Prenatal exposure to cigarette smoke and/or alcohol may contribute to abnormal fetal medullary 5-HT development in SIDS infants.

Age Factors↗

Design and analysis of trials with quality of life as an outcome: a practical guide.

Health Related Quality of Life (HRQoL) measures are becoming more frequently used in clinical trials, as both primary and secondary endpoints. Investigators are now asking statisticians for advice on how to plan (e.g., sample size) and analyze studies using HRQoL measures. HRQoL measures such as the SF-36 are usually measured on an ordered categorical (ordinal) scale. In the designing stages and when analyzing, the scales are often scored and the scores treated as if they were continuous and normally distributed. However the ordinal scaling of HRQoL measures leads to problems in determining sample size, and conventional parametric methods of estimation and hypothesis testing may not be appropriate for such outcomes. We present practical guidelines for the design and analysis of trials with HRQoL measures as outcomes. We used conventional statistical methods (i.e., t-tests and multiple regression), various ordinal regression models (proportional odds, continuation ratio, polytomous and stereotype) and bootstrap methods to analyze an HRQoL dataset. To illustrate the various methods we used HRQoL data on the SF-36 Role Limitations Emotional dimension for two groups of patients with leg ulcers. The bootstrap, t-test, and multiple regression methods gave similar results. The various ordinal regression models also gave similar results. If the HRQoL measure has a large number of ordered categories, most of which are occupied, and the underlying scale really is continuous but measured imperfectly by an instrument with a limited number of discrete values, then an informal rule of thumb is that this discrete scale should be treated as continuous if it has seven or more categories and as ordinal otherwise.

Algorithms↗