Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Victorian Emergency Minimum Dataset: factors that impact upon the data quality.

OBJECTIVE: The Victorian Emergency Minimum Dataset (VEMD) records details of approximately 80% of Victoria's ED presentations. Its usefulness for quality assurance and research relies on the data being both complete and accurate. We aimed to determine the factors that impact adversely on the collection of high-quality VEMD data. METHODS: The study was a voluntary, anonymous, cross-sectional survey of a range of ED staff (medical, nursing, clerical) who collect and enter data into the VEMD. Nine of the 28 hospitals that contribute to the VEMD were surveyed. The questionnaire was purpose-designed and self-administered. RESULTS: A total of 218 staff participated (response rate 95%). Six different software types were used, with 40% of respondents using the Pickware (MCAT) system. There was no consistency of ED personnel for the completion of specific data fields. One hundred and twenty-six (56%) respondents had heard of the VEMD, 67 (29%) had had its structure and purpose explained and 65 (30%) had been trained to enter data. Ninety-seven (45%) respondents knew what the VEMD data was used for, 38 (17%) knew they could request VEMD data for their own use and 17 (7.8%) had done so. Time constraints, software problems and lack of formal orientation and training in data entry were reported as the most important factors impacting adversely upon quality data entry. CONCLUSION: Staff knowledge of the VEMD system and its uses are poor. Numerous factors impact on the quality of data entered and interventions aimed at improving staff education, training and feedback and software are indicated.

Attitude of Health Personnel↗

Genetic heterogeneity in tuberous sclerosis. Study of a large collaborative dataset.

Tuberous sclerosis (TSC) is a multisystem autosomal dominant hamartosis whose genetics is complicated by reduced penetrance and widely varying clinical expression. Results of linkage analyses have variously suggested two different locations for a TSC gene. A collaborative dataset has been assembled to clarify the issue of genetic heterogeneity. We have now analyzed the data from a combined sample of 111 families. Using Ott's HOMOG programs, we completed three tests of homogeneity: (1) for chromosome 9q, (2) for chromosome 11q, and (3) for the combined 9q and 11q data. For test 1 the chi-square (1 df) was 21.54 (p less than 0.001), for test 2 the chi-square (1 df) was 0.13 (p greater than 0.35), and for test 3 the chi-square (2 df) was 37.61 (p less than 0.0001). Additionally, we examined the combined data for evidence that a third, as yet unlinked locus exists. Results of this last test were suggestive but not significant. Clearly loci for TSC are present on both chromosomes 9q and 11q. The maximum likelihood estimate of the proportion of chromosome 9q-linked families is 0.38, for chromosome 11q-linked families is 0.47, and for the unlinked type 0.15. Alternative explanations for these latter families include chance sampling of recombinants, nongenetic phenocopies, or misclassification.

Chromosomes, Human, Pair 11↗

Best Practice No 180. Nephrectomy for renal tumour; dissection guide and dataset.

Renal tumours constitute 2.5% of all malignancies and are among the 10 most common malignancies in the UK. Most of these are renal cell carcinomas (RCC) of various subtypes. Although historically RCC has been shown to be resistant to radiotherapy and chemotherapy, recent data suggest that the use of biological treatments, such as adjuvants, may be beneficial in patients with disease that has progressed at the time of presentation. The accurate diagnosis, staging, and grading of RCC is now a crucial element in optimal patient management. There are data to support the importance of histological type, tumour size, stage (especially patterns of extrarenal spread), and grade in determining outcome, and these data have been used to develop the published classification (Heidelberg/Rochester), staging (TNM), and grading (Fuhrman) systems. This article describes a dissection and histological sampling protocol that has been shown to increase the yield of staging information, a guide to histological classification and grading, and finally a minimum dataset for the completion of a satisfactory pathology report.

Carcinoma, Renal Cell↗

Melanoma histopathology reporting: are we complying with the National Minimum Dataset?

BACKGROUND: The Royal College of Pathologists introduced the National Minimum Dataset (NMDS) for the histopathological reporting of cutaneous melanoma in February 2002. AIM: To determine if histological reporting of invasive primary cutaneous melanoma in the West Midlands region of the UK was compliant with the NMDS. METHODS: Reports were identified from March 2002 to March 2003 via the regional Cancer Intelligence Unit, and compared with the NMDS. If all items of the NMDS were adhered to, the report was considered compliant. If not compliant, the report was checked to see if it included selected clinical and staging parameters. RESULTS: 543 cases of invasive cutaneous melanoma were identified, but only 407 reports were analysed. 69/407 (17%) (95% CI 14% to 20%) reports were fully compliant with the NMDS. Of the non-complaint reports, 45/361 (12%) (95% CI 9% to 16%) reported all staging and clinically relevant parameters; 62/361 (17%) (95% CI 59% to 65%) reported all staging parameters. Breslow thickness was reported in all but one of the reports (99.7%), Clark's level was reported in 344/407 (85%), ulceration in 280/407 (69%), and microsatellites in 146/407 (36%). CONCLUSION: There was slow uptake of the NMDS in this region in the year following its introduction. Although major parameters required for staging were more consistently reported, ulceration and microsatellites were less frequently reported.

England↗

Integrated statistical analysis of cDNA microarray and NIR spectroscopic data applied to a hemp dataset.

Both cDNA microarray and spectroscopic data provide indirect information about the chemical compounds present in the biological tissue under consideration. In this paper simple univariate and bivariate measures are used to investigate correlations between both types of high dimensional analyses. A large dataset of 42 hemp samples on which 3456 cDNA clones and 351 NIR wavelengths have been measured, was analyzed using graphical representations. For this purpose we propose clustered correlation and clustered discrimination images. Large, tissue-related differences are seen to dominate the cDNA-NIR correlation structure but smaller, more difficult to detect, variety-related differences can be found at specific cDNA clone/NIR wavelength combinations.

Algorithms↗

Predicting functional sites with an automated algorithm suitable for heterogeneous datasets.

BACKGROUND: In a previous report (La et al., Proteins, 2005), we have demonstrated that the identification of phylogenetic motifs, protein sequence fragments conserving the overall familial phylogeny, represent a promising approach for sequence/function annotation. Across a structurally and functionally heterogeneous dataset, phylogenetic motifs have been demonstrated to correspond to a wide variety of functional site archetypes, including those defined by surface loops, active site clefts, and less exposed regions. However, in our original demonstration of the technique, phylogenetic motif identification is dependent upon a manually determined similarity threshold, prohibiting large-scale application of the technique. RESULTS: In this report, we present an algorithmic approach that determines thresholds without human subjectivity. The approach relies on significant raw data preprocessing to improve signal detection. Subsequently, Partition Around Medoids Clustering (PAMC) of the similarity scores assesses sequence fragments where functional annotation remains in question. The accuracy of the approach is confirmed through comparisons to our previous (manual) results and structural analyses. Triosephosphate isomerase and arginyl-tRNA synthetase are discussed as exemplar cases. A quantitative functional site prediction assessment algorithm indicates that the phylogenetic motif predictions, which require sequence information only, are nearly as good as those from evolutionary trace methods that do incorporate structure. CONCLUSION: The automated threshold detection algorithm has been incorporated into MINER, our web-based phylogenetic motif identification server. MINER is freely available on the web at http://www.pmap.csupomona.edu/MINER/. Pre-calculated functional site predictions of the COG database and an implementation of the threshold detection algorithm, in the R statistical language, can also be accessed at the website.

Algorithms↗

CARMA: A platform for analyzing microarray datasets that incorporate replicate measures.

BACKGROUND: The incorporation of statistical models that account for experimental variability provides a necessary framework for the interpretation of microarray data. A robust experimental design coupled with an analysis of variance (ANOVA) incorporating a model that accounts for known sources of experimental variability can significantly improve the determination of differences in gene expression and estimations of their significance. RESULTS: To realize the full benefits of performing analysis of variance on microarray data we have developed CARMA, a microarray analysis platform that reads data files generated by most microarray image processing software packages, performs ANOVA using a user-defined linear model, and produces easily interpretable graphical and numeric results. No pre-processing of the data is required and user-specified parameters control most aspects of the analysis including statistical significance criterion. The software also performs location and intensity dependent lowess normalization, automatic outlier detection and removal, and accommodates missing data. CONCLUSION: CARMA provides a clear quantitative and statistical characterization of each measured gene that can be used to assess marginally acceptable measures and improve confidence in the interpretation of microarray results. Overall, applying CARMA to microarray datasets incorporating repeated measures effectively reduces the number of gene incorrectly identified as differentially expressed and results in a more robust and reliable analysis.

Analysis of Variance↗

Identification of susceptibility loci for complex diseases in a case-control association study using the Genetic Analysis Workshop 14 dataset.

Although current methods in genetic epidemiology have been extremely successful in identifying genetic loci responsible for Mendelian traits, most common diseases do not follow simple Mendelian modes of inheritance. It is important to consider how our current methodologies function in the realm of complex diseases. The aim of this study was to determine the ability of conventional association methods to fine map a locus of interest. Six study populations were selected from 10 replicates (New York) from the Genetic Analysis Workshop 14 simulated dataset and analyzed for association between the disease trait and locus D2. Genotypes from 45 single-nucleotide polymorphisms in the telomeric region of chromosome 3 were analyzed by Pearson's chi-square tests for independence to test for association with the disease trait of interest. A significant association was detected within the region; however, it was found 3 cM from the documented location of the D2 disease locus. This result was most likely due to the method used for data simulation. In general, this study showed that conventional case-control association methods could detect disease loci responsible for the development of complex traits.

Case-Control Studies↗

The effect of missing data on linkage disequilibrium mapping and haplotype association analysis in the GAW14 simulated datasets.

We used our newly developed linkage disequilibrium (LD) plotting software, JLIN, to plot linkage disequilibrium between pairs of single-nucleotide polymorphisms (SNPs) for three chromosomes of the Genetic Analysis Workshop 14 Aipotu simulated population to assess the effect of missing data on LD calculations. Our haplotype analysis program, SIMHAP, was used to assess the effect of missing data on haplotype-phenotype association. Genotype data was removed at random, at levels of 1%, 5%, and 10%, and the LD calculations and haplotype association results for these levels of missingness were compared to those for the complete dataset. It was concluded that ignoring individuals with missing data substantially affects the number of regions of LD detected which, in turn, could affect tagging SNPs chosen to generate haplotypes.

Chromosome Mapping↗

Linkage mapping of a complex trait in the New York population of the GAW14 simulated dataset: a multivariate phenotype approach.

Multivariate phenotypes underlie complex traits. Thus, instead of using the end-point trait, it may be statistically more powerful to use a multivariate phenotype correlated to the end-point trait for detecting linkage. In this study, we develop a reverse regression method to analyze linkage of Kofendrerd Personality Disorder affection status in the New York population of the Genetic Analysis Workshop 14 (GAW14) simulated dataset. When we used the multivariate phenotype, we obtained significant evidence of linkage near four of the six putative loci in at least 25% of the replicates. On the other hand, the linkage analysis based on Kofendrerd Personality Disorder status as a phenotype produced significant findings only near two of the loci and in a smaller proportion of replicates.

Chromosome Mapping↗

Critical values and variation in type I error along chromosomes in the COGA dataset using the applied pseudo-trait method.

BACKGROUND: By analyzing a "pseudo-trait," a trait not linked or associated with any of the markers tested, the distribution of the test statistic under the null hypothesis can provide the critical value for the appropriate percentile of the distribution. In addition, the anecdotal observation that p-values tend to be more significant near the telomeres was investigated. RESULTS: The applied pseudo-trait (APT) method was applied to the Affymetrix and Illumina SNPs in the Collaborative Study on the Genetics of Alcoholism dataset to determine appropriate critical values for regression of offspring on mid-parent (ROMP) and Haseman-Elston association and linkage analyses, investigating the occurrence of type I errors in different chromosomal locations, and the extent to which the critical values obtained depend on the type of pseudo-trait used. CONCLUSION: On average, the 5 percentile critical values obtained for this study were less than the expected 0.05. The distribution of p-values does not seem to depend on chromosomal position for ROMP association analysis methods, but does in some cases for Haseman-Elston linkage analysis. Results vary with different pseudo-traits.

Alcoholism↗

Identification of "pathologs" (disease-related genes) from the RIKEN mouse cDNA dataset using human curation plus FACTS, a new biological information extraction system.

BACKGROUND: A major goal in the post-genomic era is to identify and characterise disease susceptibility genes and to apply this knowledge to disease prevention and treatment. Rodents and humans have remarkably similar genomes and share closely related biochemical, physiological and pathological pathways. In this work we utilised the latest information on the mouse transcriptome as revealed by the RIKEN FANTOM2 project to identify novel human disease-related candidate genes. We define a new term "patholog" to mean a homolog of a human disease-related gene encoding a product (transcript, anti-sense or protein) potentially relevant to disease. Rather than just focus on Mendelian inheritance, we applied the analysis to all potential pathologs regardless of their inheritance pattern. RESULTS: Bioinformatic analysis and human curation of 60,770 RIKEN full-length mouse cDNA clones produced 2,578 sequences that showed similarity (70-85% identity) to known human-disease genes. Using a newly developed biological information extraction and annotation tool (FACTS) in parallel with human expert analysis of 17,051 MEDLINE scientific abstracts we identified 182 novel potential pathologs. Of these, 36 were identified by computational tools only, 49 by human expert analysis only and 97 by both methods. These pathologs were related to neoplastic (53%), hereditary (24%), immunological (5%), cardio-vascular (4%), or other (14%), disorders. CONCLUSIONS: Large scale genome projects continue to produce a vast amount of data with potential application to the study of human disease. For this potential to be realised we need intelligent strategies for data categorisation and the ability to link sequence data with relevant literature. This paper demonstrates the power of combining human expert annotation with FACTS, a newly developed bioinformatics tool, to identify novel pathologs from within large-scale mouse transcript datasets.

Animals↗

Effect of dataset selection on the topological interpretation of protein interaction networks.

BACKGROUND: Studies of the yeast protein interaction network have revealed distinct correlations between the connectivity of individual proteins within the network and the average connectivity of their neighbours. Although a number of biological mechanisms have been proposed to account for these findings, the significance and influence of the specific datasets included in these studies has not been appreciated adequately. RESULTS: We show how the use of different interaction data sets, such as those resulting from high-throughput or small-scale studies, and different modelling methodologies for the derivation pair-wise protein interactions, can dramatically change the topology of these networks. Furthermore, we show that some of the previously reported features identified in these networks may simply be the result of experimental or methodological errors and biases. CONCLUSION: When performing network-based studies, it is essential to define what is meant by the term "interaction" and this must be taken into account when interpreting the topologies of the networks generated. Consideration must be given to the type of data included and appropriate controls that take into account the idiosyncrasies of the data must be selected.

Computational Biology↗

Dataset of manually measured QT intervals in the electrocardiogram.

BACKGROUND: The QT interval and the QT dispersion are currently a subject of considerable interest. Cardiac repolarization delay is known to favor the development of arrhythmias. The QT dispersion, defined as the difference between the longest and the shortest QT intervals or as the standard deviation of the QT duration in the 12-lead ECG is assumed to be reliable predictor of cardiovascular mortality. The seventh annual PhysioNet/Computers in Cardiology Challenge, 2006 addresses a question of high clinical interest: Can the QT interval be measured by fully automated methods with accuracy acceptable for clinical evaluations? METHOD: The PTB Diagnostic ECG Database was given to 4 cardiologists and 1 biomedical engineer for manual marking of QRS onsets and T-wave ends in 458 recordings. Each recording consisted of one selected beat in lead II, chosen visually to have minimum baseline shift, noise, and artifact.In cases where no T wave could be observed or its amplitude was very small, the referees were instructed to mark a 'group-T-wave end' taking into consideration leads with better manifested T wave.A modified Delphi approach was used, which included up to three rounds of measurements to obtain results closer to the median. RESULTS: A total amount of 2*5*548 Q-onsets and T-wave ends were manually marked during round 1. To obtain closer to the median results, 8.58 % of Q-onsets and 3.21 % of the T-wave ends had to be reviewed during round 2, and 1.50 % Q-onsets and 1.17 % T-wave ends in round 3. The mean and standard deviation of the differences between the values of the referees and the median after round 3 were 2.43 +/- 0.96 ms for the Q-onset, and 7.43 +/- 3.44 ms for the T-wave end. CONCLUSION: A fully accessible, on the Internet, dataset of manually measured Q-onsets and T-wave ends was created and presented in additional file: 1 (Table 4) with this article. Thus, an available standard can be used for the development of automated methods for the detection of Q-onsets, T-wave ends and for QT interval measurements.

Algorithms↗

GOToolBox: functional analysis of gene datasets based on Gene Ontology.

We have developed methods and tools based on the Gene Ontology (GO) resource allowing the identification of statistically over- or under-represented terms in a gene dataset; the clustering of functionally related genes within a set; and the retrieval of genes sharing annotations with a query gene. GO annotations can also be constrained to a slim hierarchy or a given level of the ontology. The source codes are available upon request, and distributed under the GPL license.

Animals↗

Microbiome Datahub: an open-access platform integrating environmental metadata, taxonomy, and functional annotation for comprehensive metagenome-assembled genome datasets.

BACKGROUND: Metagenome-assembled genomes (MAGs) provide crucial insights into the genomic diversity of uncultured microbes. However, MAG datasets deposited in public repositories such as INSDC are often difficult to reuse due to heterogeneous quality, inconsistent taxonomic and functional annotations, and insufficiently curated environmental metadata. While secondary MAG databases such as MGnify, IMG/M, and SPIRE provide standardized resources, they reconstruct MAGs de novo from public metagenomic reads and therefore do not represent the original MAGs reported in publications. RESULTS: To address this gap, we developed Microbiome Datahub, an open-access platform that systematically aggregates and re-annotates original MAGs from INSDC. We collected 214,427 MAGs, predicted genes by DFAST, performed quality assessment with CheckM, standardized taxonomic assignments with GTDB-Tk, inferred 27 phenotypic traits using Bac2Feature, assigned proteins to MBGD ortholog clusters and KEGG Orthology IDs using PZLAST, and annotated environmental metadata with the Metagenome and Microbes Environmental Ontology. Across these MAGs, the average completeness was 80.5% and contamination 1.8%; notably, the most frequent values were&#x2009;>95% completeness and&#x2009;<1% contamination, indicating that the majority of MAGs are of high quality. Comparative analyses showed that Microbiome Datahub provides phylogenetically and environmentally diverse MAGs: while the majority originated from vertebrate gut environments, a substantial number were also recovered from other habitats such as groundwater, including nearly 10,000 MAGs from the Patescibacteria. Inference of 27 phenotypic traits, including optimum growth temperature, further revealed ecological differentiation across phyla. Protein clustering revealed 56 million identity 40% clusters, with the majority unique compared with MGnify and GlobDB, and&#x2009;~19% of proteins unassigned to MBGD ortholog clusters, underscoring their novelty. CONCLUSIONS: Microbiome Datahub integrates MAG genome sequences, gene and protein predictions, quality metrics, environmental and taxonomic annotations, ortholog cluster assignments, and phenotype predictions, all accessible via a web interface, API, and bulk downloads. By combining original MAGs with curated metadata and functional annotations, Microbiome Datahub constitutes a comprehensive and reusable resource that will accelerate microbiome and microbial genomics research. Video Abstract.

Metagenome↗