Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

UnionLoops: a workflow for calling chromatin loops across related Hi-C datasets with improved specificity, precision, and sensitivity.

Chromatin loop calling from chromatin interaction data often exhibits substantial variability across related samples. We present UnionLoops, a computational workflow for chromatin loop calling across multiple related samples. UnionLoops integrates information across datasets to determine positions and dataset-specificity of looping interactions. It constructs a unified candidate loop set, applies consistent filtering and aggregation, and evaluates loop support across samples. We demonstrate that UnionLoops increases sensitivity for detecting shared chromatin loops, reduces spurious sample-specific calls, and improves concordance with independent genomic features, including CTCF and cohesin occupancy. UnionLoops enables improved biological interpretation of chromatin loop organization and dynamics across related conditions.

Chromatin↗

The genesis of a catalog of oral health-related surveys: locating oral health-related datasets.

The National Institute of Dental and Craniofacial Research (NIDCR), in collaboration with the Division of Oral Health, Centers for Disease Control and Prevention (DOH, CDC), has established a Dental, Oral and Craniofacial Data Resource Center (DRC). One element of the DRC is the Catalog of Surveys Related to Oral Health. The Catalog is a searchable electronic database that includes federal, state, international, and privately sponsored surveys and other datasets. Its purpose is to make researchers aware of surveys that have been conducted and to highlight features of complex surveys that relate to oral health. Other components of the DRC include an Archive of Procedures and Methods, Archive of Procedures and Methods Used in Oral Health Surveys, which is linked to the Catalog; an Annual Report, Oral Health U.S., 2002; and a data warehouse of acquired datasets. A Web-based statistical query system related to oral health is also under development. It is the intention of the DRC to meet the needs of NIDCR and DOH, CDC staff as well as other researchers interested in the status of oral health. The Catalog is available on CD-ROM at no cost and in the future will be made available through the NIDCR Web site.

Catalogs, Library↗

Viral genome sequence datasets display pervasive evidence of strand-specific substitution biases that are best described using non-reversible nucleotide substitution models.

Most phylogenetic trees are inferred using time-reversible evolutionary models that assume that the relative rates of substitution for any given pair of nucleotides are the same regardless of the direction of the substitutions. However, there is no reason to assume that the underlying biochemical mutational processes that cause substitutions are similarly symmetrical. We consider two non-reversible nucleotide substitution models: (1) a 6-rate non-reversible model (NREV6) that is applicable to analyzing mutational processes in double-stranded genomes in that complementary substitutions occur at identical rates; and (2) a 12-rate non-reversible model (NREV12) that is applicable to analyzing mutational processes in single-stranded (ss) genomes in that all substitution types are free to occur at different rates. Using likelihood ratio and Akaike Information Criterion-based model tests, we show that, surprisingly, NREV12 provided a significantly better fit than the General Time Reversible (GTR) and NREV6 models to 21/31 dsRNA and 20/30 dsDNA datasets. As expected, however, NREV12 provided a significantly better fit to 24/33 ssDNA and 40/47 ssRNA datasets. We tested how non-reversibility impacts the accuracy with which phylogenetic trees are inferred. As simulated degrees of non-reversibility (DNR) increased, the tree topology inferences using both NREV12 and GTR became more accurate, whereas inferred tree branch lengths became less accurate. We conclude that while non-reversible models should be helpful in the analysis of mutational processes in most virus species, there is no pressing need to use these models for routine phylogenetic inference.

Models of evolution↗

SeqState: primer design and sequence statistics for phylogenetic DNA datasets.

Choosing and designing primers based on available DNA sequence data and statistical contrasting of domains or structural features is a common routine among molecular biologists. Currently available, free software tools were found to lack desirable features related to these tasks. This was the motivation for developing a new program, SeqState. SeqState locates regions that remain to be sequenced in phylogenetic DNA datasets, evaluates user-provided primers and selects primers best suited to fill gaps in the sequences. If the primers provided by the user are unsuitable, new primers are designed. Primers can be loaded from a primer database, be supplied as part of the alignment or be entered manually. The position of internal primers is automatically localised in the loaded data file. Primers can be edited, and changes and new primers can be saved to the database. Primer sheets allow the user to view internal dimers, complements to a second primer, mismatches to all loaded sequences, and other primer characteristics. Calculation of various sequence statistics can be requested for the whole dataset or parts thereof (character sets), with standard errors estimated by bootstrapping. Insertion-deletion events can be evaluated statistically and encoded for subsequent phylogenetic analysis according to several published coding principles.

Algorithms↗

The GRID: The General Repository for Interaction Datasets.

We have developed a relational database, called the General Repository for Interaction Datasets (The GRID; http://biodata.mshri.on.ca/grid) to archive and display physical, genetic and functional interactions. The GRID displays data-rich interaction tables for any protein of interest, combines literature-derived and high throughput interaction datasets, and is readily accessible via the World Wide Web. Interactions parsed in the GRID can be viewed in graphical form with a versatile visualization tool called Osprey.

Databases, Genetic↗

Small area population estimates project: data quality of administrative datasets.

The Office for National Statistics (ONS) has set up a project to investigate the feasibility of producing postcensal small area population estimates on a nationally consistent basis for England and Wales. Research has taken place to identify datasets that could potentially be used within a method to produce small area population estimates. Following an evaluation of a number of different administrative datasets, the most suitable have been short-listed for further consideration. This article presents the findings of the evaluation, based on 2001 data, and summarises the characteristics of these short-listed data sources. This article does not cover the methods that are being evaluated as part of the feasibility assessment.

Adolescent↗

Grid enabled remote visualization of medical datasets.

We present an architecture for remote visualization of datasets over the Grid. This permits an implementation-agnostic approach, where different systems can be discovered, reserved and orchestrated without being concerned about specific hardware configurations. We illustrate the utility of our approach to deliver high-quality interactive visualizations of medical datasets (circa 1 million triangles) to physically remote users, whose local physical resources would be otherwise overwhelmed. Our architecture extends to a full collaborative, resource-aware environment, whilst our presentation details our first proof-of-concept implementation.

Database Management Systems↗

A nation-wide project for the revision of the Belgian nursing minimum dataset: from concept to implementation.

This paper describes the process of revising the Belgian Nursing Minimum Data Set (NMDS). The study started in 2000. Implementation is planned from 2006. The project is divided in 4 major phases. The first phase (June-October 2002) implied the development of the conceptual framework based on literature review and secondary data-analysis. The Nursing Interventions Classification (NIC) was selected as framework for the revision of the NMDS. The second phase focused on the language development (November 2002 - September 2003) with panels of clinical experts (N=75) for six care programs. They indicated hospital financing, nurse staffing allocation and assessment of the appropriateness of hospitalization as priorities of a revised B-NMDS. A draft instrument with 84 variables, using NIC, was developed during this period. This leads to an alpha version of a revised NMDS. The third phase (October 2003 - December 2004) focused on the data collection and validation of the new tool. The new NMDS was tested on 158 nursing wards in 66 Belgian hospitals from December 2003 until March 2004. This test generated data for some 95.000 inpatient days. The interrater-reliability of the revised NMDS is tested. The criterion-related validity of the revised NMDS is compared with the actual NMDS. The discriminative power of the revised NMDS is tested to select the most relevant items for data collection. This will result in a beta version of revised NMDS in December 2004. The records of the revised NMDS are linked with the hospital discharge dataset and other mandatory datasets to integrate the revised NMDS in the broader health care management. The fourth phase (January-December 2005) will focus on information management.

Belgium↗

[Construction of standard human transcript dataset based on RefSeq and human genome sequence database].

The NCBI Reference Sequence (RefSeq) database aimed to provide a biologically non-redundant collection of DNA, RNA, and protein sequences and to promote the research on genes and proteins of human beings and other species. However, because of widely distributed polymorphisms and different quality control of experiments in individual laboratories, there are potential problems need to be identified in the RefSeq database. Regarding which, we herein define the concept, standard transcript, based on the Central Dogmas of Biology that each standard transcript should be perfectly mapped to the standard genomic DNA sequence at the exon level. A large scale analysis for mapping all of the RefSeq records of human being (2005-4-18) to the officially released human genome sequence database (2005-4-20) was further performed using BLAT, Sim4 and a homemade program, EIparser, which was especially designed for this purpose. The standard transcripts based on the RefSeq database were obtained according to the alignment with standard human genome database. There are 9,771 RefSeq records of human being labeled with "NM_" and "NR_" could be perfectly mapped to human genome sequences, while other 10,943 records could be considered as standard transcripts after reasonable revision by comparing with the genome sequences according to all of the three methods. Moreover, the left 203 unrevisable records and 2,676 inconsistent records reported by the above programs could not be considered as standard transcripts and should be checked critically before using because of potential errors in them. Our study has thus provided a reference standard dataset of human beings with high quality for further bioinformatic and experimental analysis such as polymorphism and mutation of human genes. The reference standard dataset based on above criteria could be retrieved from http://biocompute.bmi.ac.cn/transcriptome/index.htm.

Databases, Genetic↗

Reviewing and managing syndromic surveillance SaTScan datasets using an open source data visualization tool.

SaTScan is a popular, free software tool used to identify disease clusters early in the course of an outbreak. Using geographic and time-based surveillance data, SaTScan can generate large datasets that are difficult for humans to interpret. Tracing disease clusters through space and time using text tables is a challenging cognitive task. To simplify this process, we developed a Java-based open-source tool to transform SaTScan analytic datasets into easily navigable data visualizations.

Cluster Analysis↗

Developing a data dictionary for the irish nursing minimum dataset.

One of the challenges in health care in Ireland is the relatively slow acceptance of standardised clinical information systems. Yet the national Irish health reform programme indicates that an Electronic Health Care Record (EHCR) will be implemented on a phased basis. [3-5]. While nursing has a key role in ensuring the quality and comparability of health information, the so- called 'invisibility' of some nursing activities makes this a challenging aim to achieve [3-5]. Any integrated health care system requires the adoption of uniform standards for electronic data exchange [1-2]. One of the pre-requisites for uniform standards is the composition of a data dictionary. Inadequate definition of data elements in a particular dataset hinders the development of an integrated data depository or electronic health care record (EHCR). This paper outlines how work on the data dictionary for the Irish Nursing Minimum Dataset (INMDS) has addressed this issue. Data set elements were devised on the basis of a large scale empirical research programme. ISO 18104, the reference terminology for nursing [6], was used to cross-map the data set elements with semantic domains, categories and links and data set items were dissected.

Databases as Topic↗

[EDP registration of the types of tests, working procedures and diagnosis in a department of histopathology. Minimum dataset for assessment of the workload].

The aim of the study was to develop a computer-based system for measuring workload in a department of histo- and cytopathology using routine registration of a minimum dataset. A group with representatives from the laboratory technicians, the pathologists and the secretaries defined 18 types of specimens. By studying each step of specimen processing it was shown, that 14 items for technical details could cover all the work done in the department. This information was collected in a computer-based system connected to the hospital network. The measurement of workload is essential for the efficient management of laboratory services. The registration of specimen types and a minimum dataset for specimen processing describes the work done in a department of histo- and cytopathology.

Computer Systems↗

Integrated genomic analysis of NF1-associated peripheral nerve sheath tumors: an updated biorepository dataset.

Neurofibromatosis type 1 (NF1) is an inherited neurocutaneous condition that predisposes to the development of peripheral nerve sheath tumors (PNST) including cutaneous neurofibromas (CNF), plexiform neurofibromas (PNF), atypical neurofibromatous neoplasms of uncertain biologic potential (ANNUBP), and malignant peripheral nerve sheath tumors (MPNST). The Johns Hopkins NF1 biospecimen repository promotes the successful advancement of therapeutic developments for NF1-associated PNST through acquisition and genomic analysis of human tumor specimens. RNA sequencing (RNAseq) and whole exome sequencing (WES) data were generated from 73 and 114 primary human tumor samples, respectively. These pre-processed data, standardized for immediate computational analysis, are accessible through the NF Data Portal, allowing immediate interrogation. This dataset combines new and previously released samples, offering a comprehensive view of the entire cohort sequenced. As a dedicated effort to systematically bank tumor samples from people with NF1, in collaboration with molecular geneticists and computational biologists, the Johns Hopkins NF1 biospecimen repository offers access to tissue samples and genomic data to promote the advancement of NF1-related tumor biologic insights and therapies.

Humans↗

16S rRNA and Metagenomic Datasets of Gastrointestinal Microbiota in Fetal and 7-Day-Old Goat Kids.

The perinatal period (from late gestation to the neonatal stage) in ruminants is a critical phase for fetal organ maturation, where ecological succession of gastrointestinal microbial communities significantly impacts livestock production efficiency. However, research remains insufficient regarding the distribution patterns and functional annotation of microbial communities across different gastrointestinal compartments during this period. This study characterized early microbiota dynamics in Hutianshi Goats using 16S rRNA sequencing (4 fetal goats at 90 ± 10 gestational days) and metagenomics (3 7-day-old goat kids). The fetal goat group generated 852,694 valid reads, yielding 688,277 high-quality reads after chimera removal for downstream analysis. The 7-day-old goat kids group produced 1,081,588,182 final valid reads, after data processing and assembly, 8,561,345 contigs were generated. Gene prediction identified 6,095,352 genes. Multi-database annotations (NR, KEGG, CAZy, etc.) revealed functional potential and antimicrobial resistance traits. The public release of this dataset facilitates academic understanding of microbial community dynamics and host-microbe interactions during this developmental stage, providing both theoretical foundations and data resources for ruminant developmental biology and precision breeding regulation.

Animals↗

Metagenomic and Transcriptomic Datasets of Plateau Brown Frogs (Rana kukunoris) from the Helan Mountains.

Global climate change has become a primary driving factor behind the biodiversity crisis in amphibians, making it crucial to understand how climate change affects species and their potential responses. The plateau brown frog (Rana kukunoris) is often regarded as an ideal ecological indicator species, yet research on its environmental adaptation mechanisms based on transcriptomic and microbiomic studies remains limited. Therefore, this study investigates the adaptation strategies of the plateau brown frog to environmental changes, providing extensive transcriptomic and the first comprehensive metagenomic dataset from two distinctly different environmental regions (eastern and western slopes of the Helan Mountains). We gathered transcriptomic data from three tissues (blood, liver, and muscle), resulting in 294,962 unigenes and 570,192 transcripts. Metagenomic sequencing identified major bacterial groups, including Firmicutes, Proteobacteria, Bacteroidetes, Spirochetes, and Actinobacteria. In summary, the results of this study can be used to further explore the associations among microbiota, host, and environment, which are crucial for comprehending the mechanisms of environmental adaptation in this species and contributing to the conservation of amphibian biodiversity.

Animals↗

Some conceptual and statistical issues in analysis of longitudinal psychiatric data. Application to the NIMH treatment of Depression Collaborative Research Program dataset.

Longitudinal studies have a prominent role in psychiatric research; however, statistical methods for analyzing these data are rarely commensurate with the effort involved in their acquisition. Frequently the majority of data are discarded and a simple end-point analysis is performed. In other cases, so called repeated-measures analysis of variance procedures are used with little regard to their restrictive and often unrealistic assumptions and the effect of missing data on the statistical properties of their estimates. We explored the unique features of longitudinal psychiatric data from both statistical and conceptual perspectives. We used a family of statistical models termed random regression models that provide a more realistic approach to analysis of longitudinal psychiatric data. Random regression models provide solutions to commonly observed problems of missing data, serial correlation, time-varying covariates, and irregular measurement occasions, and they accommodate systematic person-specific deviations from the average time trend. Properties of these models were compared with traditional approaches at a conceptual level. The approach was then illustrated in a new analysis of the National Institute of Mental Health Treatment of Depression Collaborative Research Program dataset, which investigated two forms of psychotherapy, pharmacotherapy with clinical management, and a placebo with clinical management control. Results indicated that both person-specific effects and serial correlation play major roles in the longitudinal psychiatric response process. Ignoring either of these effects produces misleading estimates of uncertainty that form the basis of statistical tests of hypotheses.

Analysis of Variance↗

Biomedical image visualization research using the Visible Human Datasets.

The practice of medicine and conduct of research in major segments of the biologic sciences have always relied on visualizations to study the relationship of anatomic structure to biologic function. Traditionally, these visualizations have either been direct, via vivisection and postmortem examination, or have required extensive mental reconstruction. The revolutionary capabilities of 3-D and 4-D medical imaging modalities, together with computer reconstruction and rendering of multidimensional medical and histological volume image data, obviate the need for physical dissection or abstract assembly. The availability of the Visible Human Datasets from the National Library of Medicine, coupled with the development of advanced computer algorithms to accurately and rapidly process, segment, register, measure, and display high resolution 3-D images, has provided a rich opportunity to help advance these important new imaging, visualization, and analysis methodologies from scientific theory to clinical practice.

Algorithms↗

Exploring predictive and reproducible modeling with the single-subject FIAC dataset.

Predictive modeling of functional magnetic resonance imaging (fMRI) has the potential to expand the amount of information extracted and to enhance our understanding of brain systems by predicting brain states, rather than emphasizing the standard spatial mapping. Based on the block datasets of Functional Imaging Analysis Contest (FIAC) Subject 3, we demonstrate the potential and pitfalls of predictive modeling in fMRI analysis by investigating the performance of five models (linear discriminant analysis, logistic regression, linear support vector machine, Gaussian naive Bayes, and a variant) as a function of preprocessing steps and feature selection methods. We found that: (1) independent of the model, temporal detrending and feature selection assisted in building a more accurate predictive model; (2) the linear support vector machine and logistic regression often performed better than either of the Gaussian naive Bayes models in terms of the optimal prediction accuracy; and (3) the optimal prediction accuracy obtained in a feature space using principal components was typically lower than that obtained in a voxel space, given the same model and same preprocessing. We show that due to the existence of artifacts from different sources, high prediction accuracy alone does not guarantee that a classifier is learning a pattern of brain activity that might be usefully visualized, although cross-validation methods do provide fairly unbiased estimates of true prediction accuracy. The trade-off between the prediction accuracy and the reproducibility of the spatial pattern should be carefully considered in predictive modeling of fMRI. We suggest that unless the experimental goal is brain-state classification of new scans on well-defined spatial features, prediction alone should not be used as an optimization procedure in fMRI data analysis.

Artifacts↗