Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Datasets as Topic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

Tetanus toxoid and congenital abnormalities.

OBJECTIVE: To study the human teratogenic potential of tetanus vaccination during pregnancy. METHODS: Pair analysis of cases with congenital abnormalities and matched healthy controls was performed in the large population-based dataset of the Hungarian Case-Control Surveillance of Congenital Abnormalities, 1980-1994. RESULTS: Of 35 727 pregnant women who had babies without any defects in the study period (control group), 33 (0.09%) were vaccinated with tetanus. Of 21563 pregnant women who had offspring with congenital abnormalities, 25 (0.12%) had tetanus vaccination. This difference was not significant (P = 0.39). The case-control pair analysis confirmed the safety of tetanus vaccination during pregnancy, particularly in the second and third months of gestation, i.e. during the critical period for congenital abnormalities. CONCLUSION: Tetanus vaccination during pregnancy appears not be teratogenic to the fetus. Thus, there is no contraindication, if the use of tetanus toxoid is necessary during pregnancy.

Abnormalities, Drug-Induced↗

Employing a composite gene-flow index to numerically quantify a crop's potential for gene flow: an Irish perspective.

Guidelines to ensure the efficient coexistence of genetically modified (GM) and conventional crops are currently being considered across the European Union. The purpose of this strategy is to describe the measures a farmer must adopt to minimize the admixture of GM and non-GM crops. Minimizing pollen/seed-mediated gene flow between GM and non-GM crops is central to successful coexistence. However no system is currently available to permit the numeric quantification of a crop's propensity for pollen/seed-mediated gene flow. The provision of such a system could permit a background level of gene flow, specific for a particular conventional crop, to be calculated. Here we present a gene flow index model implemented using the principal arable crops in Ireland as a model dataset. The objective of this research was to establish a baseline gene flow data set for Ireland's primary conventional crops through the provision of a simple numerical index. This Gene Flow Index (GFI) incorporates four strands of crop-mediated gene flow (crop pollen-to-crop, crop pollen-to-wild, crop seed-to-volunteer and crop seed-to-feral) into a format that permits the calculation of a crop's gene flow potential. Responsive to regional parameters, we have applied the model to sugar beet, oilseed rape, potato, ryegrass, maize, wheat and barley. We propose that the attained indices will highlight those crops that require additional measures in order to minimize gene flow in accordance with anticipated coexistence guidelines.

Crops, Agricultural↗

Application of fast Fourier transform cross-correlation for the alignment of large chromatographic and spectral datasets.

Preprocessing of chromatographic and spectral data is an important aspect of analytical sciences. In particular, recent advances in proteomics have resulted in the generation of large data sets that require analysis. To assist accurate comparison of chemical signals, we propose two methods for the alignment of multiple spectral data sets. Based on methods previously described, each chromatograph or spectrum to be aligned is divided and aligned as individual segments to a reference. However, our methods make use of fast Fourier transform for the rapid computation of a cross-correlation function that enables alignments between samples to be optimized. The proposed methods are demonstrated in comparison with an existing method on a chromatographic and a mass spectral data set. It is shown that our methods provide an advantage of speed and a reduction of the number of input parameters required. The software implementations for the proposed alignment methods are available under the downloads section at http://ptcl.chem.ox.ac.uk/~jwong/specalign.

Algorithms↗

An E-health solution for automatic sleep classification according to Rechtschaffen and Kales: validation study of the Somnolyzer 24 x 7 utilizing the Siesta database.

To date, the only standard for the classification of sleep-EEG recordings that has found worldwide acceptance are the rules published in 1968 by Rechtschaffen and Kales. Even though several attempts have been made to automate the classification process, so far no method has been published that has proven its validity in a study including a sufficiently large number of controls and patients of all adult age ranges. The present paper describes the development and optimization of an automatic classification system that is based on one central EEG channel, two EOG channels and one chin EMG channel. It adheres to the decision rules for visual scoring as closely as possible and includes a structured quality control procedure by a human expert. The final system (Somnolyzer 24 x 7) consists of a raw data quality check, a feature extraction algorithm (density and intensity of sleep/wake-related patterns such as sleep spindles, delta waves, SEMs and REMs), a feature matrix plausibility check, a classifier designed as an expert system, a rule-based smoothing procedure for the start and the end of stages REM, and finally a statistical comparison to age- and sex-matched normal healthy controls (Siesta Spot Report). The expert system considers different prior probabilities of stage changes depending on the preceding sleep stage, the occurrence of a movement arousal and the position of the epoch within the NREM/REM sleep cycles. Moreover, results obtained with and without using the chin EMG signal are combined. The Siesta polysomnographic database (590 recordings in both normal healthy subjects aged 20-95 years and patients suffering from organic or nonorganic sleep disorders) was split into two halves, which were randomly assigned to a training and a validation set, respectively. The final validation revealed an overall epoch-by-epoch agreement of 80% (Cohen's kappa: 0.72) between the Somnolyzer 24 x 7 and the human expert scoring, as compared with an inter-rater reliability of 77% (Cohen's kappa: 0.68) between two human experts scoring the same dataset. Two Somnolyzer 24 x 7 analyses (including a structured quality control by two human experts) revealed an inter-rater reliability close to 1 (Cohen's kappa: 0.991), which confirmed that the variability induced by the quality control procedure, whereby approximately 1% of the epochs (in 9.5% of the recordings) are changed, can definitely be neglected. Thus, the validation study proved the high reliability and validity of the Somnolyzer 24 x 7 and demonstrated its applicability in clinical routine and sleep studies.

Adult↗

Scientific and cost-effectiveness criteria in selecting batteries of short-term tests.

The scientific and cost-effectiveness criteria introduced in this paper can be applied to published datasets and current and proposed batteries of short-term tests. The reports in the current volume will provide a wealth of additional material for such evaluations, but more systematically obtained information will be necessary to assess both the internal and external validity of these tests. Individual tests and batteries of tests should be standardized, employ positive controls, generate results capable of quantitative analyses that may make dichotomous classification as "positive" and "negative" obsolete, be interpreted in light of mechanisms of action, and be cost-effective on a grand scale. For regulatory purposes our long-term goal should be to replace the whole animal lifetime bioassay with an appropriate and cost-effective set of short-term tests.

Animals↗

Distributed data analysis in a multicenter study: the CARDIA Study.

Unlike distributed data entry, which is used in many large epidemiologic studies and multicenter clinical trials, distributed data analysis is a relatively new concept. This paper reports on the usefulness of such a system in the Coronary Artery Risk Development in Young Adults (CARDIA) Study. CARDIA distributes the entire examination dataset to participating centers soon after completion of each round of data collection. The process was designed to encourage more numerous, diverse, and rapid publications, and to allow for more efficient use of the manpower and expertise in centers. Responsibilities of the coordinating center have changed from a conventional coordinating center but remain substantial due to the need for collating, monitoring, verifying, and documenting the distributed data analysis (DDA) system. DDA is successful from the standpoint of implementation and operation--21 manuscripts representing work analyzed at six participating centers had been submitted for publication within 3.5 years of the completion of the baseline examination.

Adolescent↗

Application of classification-tree models to characterize the mycobiota of grapes on the basis of origin.

Classification-tree (CT) models are a simple and robust exploratory data analysis technique that can be used in classification, regressions and summaries of data. They distill complex ecological relationships into simplified rules and identify the species necessary for sample classification on the basis of detailed ecological inventories. The usefulness of this technique to characterize and represent differences in the grape mycobiota of distinct origins was evaluated. Grapes from four Portuguese winemaking regions were selected for a 3-year study: Alentejo, Douro, Ribatejo and Vinhos Verdes. The mycobiota of grapes was assessed with plating methods and the frequencies of isolations of the fungal taxa identified in 32 samples were used as a training dataset. The CT algorithm selected the fungal taxa and respective frequencies to classify grapes according to its region of origin. The ten-fold cross-validation technique was used for model evaluation. The success rate of the model was quantified and expressed in the number of correctly classified samples overall and into region. Furthermore, model refinement was performed using attribute selection algorithms and class redefinition. A simple tree model was generated that classified grapes into three regional origins: Douro, South (Alentejo and Ribatejo classes together) and Vinhos Verdes, on the basis of the incidence of Aspergillus niger aggregate and Penicillium thomii in samples with an accuracy of 82%. The merits and demerits of these models are discussed.

Algorithms↗

The informatics of a C57BL/6J mouse brain atlas.

The Mouse Atlas Project (MAP) aims to produce a framework for organizing and analyzing the large volumes of neuroscientific data produced by the proliferation of genetically modified animals. Atlases provide an invaluable aid in understanding the impact of genetic manipulations by providing a standard for comparison. We use a digital atlas as the hub of an informatics network, correlating imaging data, such as structural imaging and histology, with text-based data, such as nomenclature, connections, and references. We generated brain volumes using magnetic resonance microscopy (MRM), classical histology, and immunohistochemistry, and registered them into a common and defined coordinate system. Specially designed viewers were developed in order to visualize multiple datasets simultaneously and to coordinate between textual and image data. Researchers can navigate through the brain interchangeably, in either a text-based or image-based representation that automatically updates information as they move. The atlas also allows the independent entry of other types of data, the facile retrieval of information, and the straight-forward display of images. In conjunction with centralized servers, image and text data can be kept current and can decrease the burden on individual researchers' computers. A comprehensive framework that encompasses many forms of information in the context of anatomic imaging holds tremendous promise for producing new insights. The atlas and associated tools can be found at http://www.loni.ucla.edu/MAP.

Anatomy, Artistic↗

Seamless multiresolution display of portable wavelet-compressed images.

Image storage, display, and distribution have been difficult problems in radiology for many years. As improvements in technology have changed the nature of the storage and display media, demand for image portability, faster image acquisition, and flexible image distribution is driving the development of responsive systems. Technology, such as the wavelet-based multiresolution seamless image database (MrSID) portable image format (PIF), is enabling image management solutions that address the shifting "point-of-care." The MrSID PIF employs seamless, multiresolution technology, which allows the viewer to determine the size of the image to be viewed, as well as the position of the viewing area within the image dataset. In addition the MrSID PIF allows control of the compression ratio of decompressed images. This capability offers the advantage of very rapid image recall from storage devices and portability for rapid transmission and distribution using the internet or wide-area networks. For example, in teleradiology, the radiologist or other physician desiring to view images at a remote location has full flexibility in being able to choose a quick display of an overview image, a complete display of a full diagnostic quality image, or both without compromising communication bandwidth. The MrSID algorithm will satisfy Joint Photographic Experts Group (JPEG) 2000 standards, thereby being compatible with future versions of the Digital Imaging and Communications in Medicine (DICOM) standard for image data compression.

Algorithms↗

Discovering common stem-loop motifs in unaligned RNA sequences.

Post-transcriptional regulation of gene expression is often accomplished by proteins binding to specific sequence motifs in mRNA molecules, to affect their translation or stability. The motifs are often composed of a combination of sequence and structural constraints such that the overall structure is preserved even though much of the primary sequence is variable. While several methods exist to discover transcriptional regulatory sites in the DNA sequences of coregulated genes, the RNA motif discovery problem is much more difficult because of covariation in the positions. We describe the combined use of two approaches for RNA structure prediction, FOLDALIGN and COVE, that together can discover and model stem-loop RNA motifs in unaligned sequences, such as UTRs from post-transcriptionally coregulated genes. We evaluate the method on two datasets, one a section of rRNA genes with randomly truncated ends so that a global alignment is not possible, and the other a hyper-variable collection of IRE-like elements that were inserted into randomized UTR sequences. In both cases the combined method identified the motifs correctly, and in the rRNA example we show that it is capable of determining the structure, which includes bulge and internal loops as well as a variable length hairpin loop. Those automated results are quantitatively evaluated and found to agree closely with structures contained in curated databases, with correlation coefficients up to 0.9. A basic server, Stem-Loop Align SearcH (SLASH), which will perform stem-loop searches in unaligned RNA sequences, is available at http://www.bioinf.au.dk/slash/.

Algorithms↗

Commentary and opinion: III. Some nonontological and functionally unconnected views on current issues in the analysis of PET datasets.

Strother et al. (1995) and Friston (1995) both raise important issues and provide useful reviews of various aspects of PET data analysis. Statisticians would not assume that any single piece of methodology would answer all questions about a type of data in a variety of experimental and observational contexts. The fundamental importance of hypothesis-driven inference, based on well designed experiments, cannot be overestimated for its ability to progress scientific understanding in an orderly manner. However, hypothesis-generating experiments are also vital in their own right. In practice, we generally do not have the luxury of both types of experiment, and we should note Strother et al.'s comment on the importance of extracting as much information as possible from each dataset. Friston (1995) also sees formal testing methods and exploratory methods such as principal components analysis as complementary. The correct approach would therefore seem to be (a) to select methods for formal and exploratory data analysis from the rich existing tool kit of statistical procedures, (b) to modify these as necessary to deal with special PET problems such as multiplicity, (c) to be aware of the assumptions underlying the methods being used and to investigate the problems that can arise if these assumptions fail to hold, (d) to appreciate the complexity of both PET data and of the potential questions that can be asked of it, and (e) to be aware of the limitations of any statistical analysis and the need for caution in interpreting conclusions not based on any predefined hypothesis.

Analysis of Variance↗

HIV cohort collaborations: proposal for harmonization of data exchange.

HIV cohort studies have provided useful information on the natural history of HIV infection and the effects of antiretroviral therapy. It has become increasingly common to combine data from several cohorts into one dataset in order to address certain specific questions with more statistical power than can be achieved with the individual studies. This requires each cohort to map data into a standard format before merging. Until recently, this standard format has differed for each such collaborative analysis. We have therefore developed the HIV Cohort Data Exchange Protocol (HICDEP), which is freely available at http://www.cphiv.dk/HICDEP.pdf. Once individual cohorts have set up a means of transfering data into this format, as and when required, this should greatly facilitate data merging for future joint analyses. The HICDEP incorporates data from HIV drug resistance tests, which have been particularly challenging for cohorts to integrate into databases.

Adverse Drug Reaction Reporting Systems↗

Establishing an international reference image database for research and development in medical image processing.

INTRODUCTION: The lack of comparability of evaluation results is one of the major obstacles of research and development in Medical Image Processing (MIP). The main reason for that is the usage of different image datasets with different quality, size and Gold standard. OBJECTIVES: Therefore, one of the goals of the Working Group on Medical Image Processing of the European Federation for Medical Informatics (EFMI WG MIP) is to develop first parts of a Reference Image Database. METHODS: Kernel of the concept is to identify highly relevant medical problems with significant potential for improvement by MIP, and then to provide respective reference datasets. The EFMI WG MIP has primarily the role of a specifying group and an information broker, while the provider user relationships are defined by bilateral co-operation or license agreements. RESULTS: An explorative database prototype has been implemented using the MySQL database software on the Web. Templates for provider user agreements have been worked out and already applied for own 'pre-RID-MIP' co-operations of the authors. DISCUSSION AND CONCLUSION: First steps towards a comprehensive reference image database have been done. Issues like funding, motivation, management, provision of Gold standards and evaluation guidelines are to be solved. Due to the interest from research groups and industry the efforts will be continued.

Databases as Topic↗

Risk factors in carpal tunnel syndrome.

We have undertaken a large case-control study using the UK General Practice Research Database to quantify the relative contributions of the common risk factors for carpal tunnel syndrome (CTS) in the community. Cases were patients with a diagnosis of CTS and, for each, four controls were individually matched by age, sex and general practice. Our dataset included 3,391 cases, of which 2,444 (72%) were women, with a mean age at diagnosis of 46 (range 16-96) years. Multivariate analysis showed that the risk factors associated with CTS were previous wrist fracture (OR=2.29), rheumatoid arthritis (OR=2.23), osteoarthritis of the wrist and carpus (OR=1.89), obesity (OR=2.06), diabetes (OR=1.51), and the use of insulin (OR=1.52), sulphonylureas (OR=1.45), metformin (OR=1.20) and thyroxine (OR=1.36). Smoking, hormone replacement therapy, the combined oral contraceptive pill and oral corticosteroids were not associated with CTS. The results were similar when cases were restricted to those who had undergone carpal tunnel decompression.

Adolescent↗

Statistical analysis of intrahelical ionic interactions in alpha-helices and coiled coils.

There are many controversies concerning whether ionic interactions in alpha-helices and coiled coils actually contribute to the stabilisation and formation of these structures. Here we used a statistical approach to probe this question. We extracted unique alpha-helical and coiled coil structures from the protein database and analysed the ionic interactions between positively and negatively charged residues. The ionic interactions were categorized according to the type, spacing and order of the residues involved. Separate datasets were produced depending on the number of alpha-helices in the coiled coils and the mutual orientation of the helices. We compared the frequency of residue configurations able to form ionic interactions with their probability to form the interaction. We found a correlation between the two variables in alpha-helices, antiparallel two-stranded coiled coils and parallel two-stranded coiled coils. This indicates that some ionic interactions are indeed important for the formation and stabilisation of alpha-helices and coiled coils. We concluded that the configurations, which have simultaneously a large probability to form the ionic interaction and a frequent occurrence, are those, which have the most stabilising effect. These are the 4RE, 3ER and 4ER interactions.

Algorithms↗

Phylogenetic relationships within mammalian order Carnivora indicated by sequences of two nuclear DNA genes.

Phylogenetic relationships among 37 living species of order Carnivora spanning a relatively broad range of divergence times and taxonomic levels were examined using nuclear sequence data from exon 1 of the IRBP gene (approximately 1.3 kb) and first intron of the TTR gene (approximately 1 kb). These data were used to analyze carnivoran phylogeny at the family and generic level as well as the interspecific relationships within recently derived Felidae. Phylogenetic results using a combined IRBP+TTR dataset strongly supported within the superfamily Califormia, the red panda as the closest lineage to procyonid-mustelid (i.e., Musteloidea) clade followed by pinnipeds (Otariidae and Phocidae), Ursidae (including the giant panda), and Canidae. Four feliform families, namely the monophyletic Herpestidae, Hyaenidae, and Felidae, as well as the paraphyletic Viverridae were consistently recovered convincingly. The utilities of these two gene segments for the phylogenetic analyses were extensively explored and both were found to be fairly informative for higher-group associations within the order Carnivora, but not for those of low level divergence at the species level. Therefore, there is a need to find additional genetic markers with more rapid mutation rates that would be diagnostic at deciphering relatively recent relationships within the Carnivora.

Animals↗

Analysis of penetrance and expressivity during ontogenesis supports a stochastic choice of zebrafish odorant receptors from predetermined groups of receptor genes.

Olfactory receptor neurons select a single odourant receptor gene for expression out of a large gene family. The mechanisms of this extreme selectivity are largely unknown. We have determined in detail the developmental expression dynamics of a representative subset of the zebrafish odourant receptor repertoire, using in situ hybridization analysis. We have thus generated a dataset, which allows us to test hypotheses of odourant receptor gene regulation. The receptors chosen belong to four different groups with respect to ontogenetic onset of expression (onset groups). Statistical analysis of the data supports a model in which the final choice of an individual odourant receptor gene occurs stochastically from within a group of genes sharing a deterministically defined onset of expression. Genomic mapping revealed a pronounced correlation of onset of expression with genomic neighbourhood. During a protracted juvenile developmental period individual regulatory influences seem to modify the expression of odourant receptor genes, a notable example being a transient decrease in expressivity of two odourant receptor genes.

Animals↗

A comparison of infection control software for use by hospital epidemiologists in meeting the new JCAHO standards.

To choose a microcomputer software package for our hospital epidemiology division, the two leading commercial software packages for infection control, AICE (ICPA, Inc., Austin, Texas) and NOS0-3 (Epi Systematics, Inc., Ft. Meyers, Florida), were compared for the types of epidemiologic analysis likely to be required to satisfy new Joint Commission on Accreditation of Healthcare Organizations (JCAHO) 1990 Infection Control Standards. The test dataset was a surgical database of 3,235 operations with 292 (9%) wound infections. Though NOSO-3 was more flexible in terms of the amount of data items one could record, it required seven times longer to learn, nine times more disk space to store and two times as long to enter cases than AICE. Six simple infection control reports (i.e., line listings, crosstabulations, stratified rates and graphs) required only seven computing steps and approximately 11 minutes to process with AICE, but 22 steps and over two hours with NOSO-3. All analytic results from AICE agreed with the results obtained with the Statistical Analysis System (SAS, SAS Institute, Inc., Cary, North Carolina), but analyses such as service-specific rates performed with NOSO-3 differed because of a design flaw in the NOSO-3 data structure.

Cross Infection↗