Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

RRE: a tool for the extraction of non-coding regions surrounding annotated genes from genomic datasets.

UNLABELLED: RRE allows the extraction of non-coding regions surrounding a coding sequence [i.e. gene upstream region, 5'-untranslated region (5'-UTR), introns, 3'-UTR, downstream region] from annotated genomic datasets available at NCBI. AVAILABILITY: RRE parser and web-based interface are accessible at http://www.bioinformatica.unito.it/bioinformatics/rre/rre.html

Chromosome Mapping↗

MIPS bacterial genomes functional annotation benchmark dataset.

MOTIVATION: Any development of new methods for automatic functional annotation of proteins according to their sequences requires high-quality data (as benchmark) as well as tedious preparatory work to generate sequence parameters required as input data for the machine learning methods. Different program settings and incompatible protocols make a comparison of the analyzed methods difficult. RESULTS: The MIPS Bacterial Functional Annotation Benchmark dataset (MIPS-BFAB) is a new, high-quality resource comprising four bacterial genomes manually annotated according to the MIPS functional catalogue (FunCat). These resources include precalculated sequence parameters, such as sequence similarity scores, InterPro domain composition and other parameters that could be used to develop and benchmark methods for functional annotation of bacterial protein sequences. These data are provided in XML format and can be used by scientists who are not necessarily experts in genome annotation. AVAILABILITY: BFAB is available at http://mips.gsf.de/proj/bfab

Bacterial Proteins↗

Avoiding model selection bias in small-sample genomic datasets.

MOTIVATION: Genomic datasets generated by high-throughput technologies are typically characterized by a moderate number of samples and a large number of measurements per sample. As a consequence, classification models are commonly compared based on resampling techniques. This investigation discusses the conceptual difficulties involved in comparative classification studies. Conclusions derived from such studies are often optimistically biased, because the apparent differences in performance are usually not controlled in a statistically stringent framework taking into account the adopted sampling strategy. We investigate this problem by means of a comparison of various classifiers in the context of multiclass microarray data. RESULTS: Commonly used accuracy-based performance values, with or without confidence intervals, are inadequate for comparing classifiers for small-sample data. We present a statistical methodology that avoids bias in cross-validated model selection in the context of small-sample scenarios. This methodology is valid for both k-fold cross-validation and repeated random sampling.

Algorithms↗

AClAP, Autonomous hierarchical agglomerative Cluster Analysis based protocol to partition conformational datasets.

MOTIVATION: Sampling the conformational space is a fundamental step for both ligand- and structure-based drug design. However, the rational organization of different molecular conformations still remains a challenge. In fact, for drug design applications, the sampling process provides a redundant conformation set whose thorough analysis can be intensive, or even prohibitive. We propose a statistical approach based on cluster analysis aimed at rationalizing the output of methods such as Monte Carlo, genetic, and reconstruction algorithms. Although some software already implements clustering procedures, at present, a universally accepted protocol is still missing. RESULTS: We integrated hierarchical agglomerative cluster analysis with a clusterability assessment method and a user independent cutting rule, to form a global protocol that we implemented in a MATLAB metalanguage program (AClAP). We tested it on the conformational space of a quite diverse set of drugs generated via Metropolis Monte Carlo simulation, and on the poses we obtained by reiterated docking runs performed by four widespread programs. In our tests, AClAP proved to remarkably reduce the dimensionality of the original datasets at a negligible computational cost. Moreover, when applied to the outcomes of many docking programs together, it was able to point to the crystallographic pose. AVAILABILITY: AClAP is available at the "AClAP" section of the website http://www.scfarm.unibo.it.

Algorithms↗

Integrating transcription factor binding site information with gene expression datasets.

MOTIVATION: Microarrays are widely used to measure gene expression differences between sets of biological samples. Many of these differences will be due to differences in the activities of transcription factors. In principle, these differences can be detected by associating motifs in promoters with differences in gene expression levels between the groups. In practice, this is hard to do. RESULTS: We combine correspondence analysis, between group analysis and co-inertia analysis to determine which motifs, from a database of promoter motifs, are strongly associated with differences in gene expression levels. Given a database of motifs and gene expression levels from a set of arrays, the method produces a ranked list of motifs associated with any specified split in the arrays. We give an example using the Gene Atlas compendium of gene expression levels for human tissues where we search for motifs that are associated with expression in central nervous system (CNS) or muscle tissues. Most of the motifs that we find are known from previous work to be strongly associated with expression in CNS or muscle. We give a second example using a published prostate cancer dataset where we can simply and clearly find which transcriptional pathways are associated with differences between benign and metastatic samples. AVAILABILITY: The source code is freely available upon request from the authors.

Algorithms↗

A tail strength measure for assessing the overall univariate significance in a dataset.

We propose an overall measure of significance for a set of hypothesis tests. The 'tail strength' is a simple function of the p-values computed for each of the tests. This measure is useful, for example, in assessing the overall univariate strength of a large set of features in microarray and other genomic and biomedical studies. It also has a simple relationship to the false discovery rate of the collection of tests. We derive the asymptotic distribution of the tail strength measure, and illustrate its use on a number of real datasets.

Biometry↗

Health education and promotion spending in England: a note on the potential utility of the Health Service Indicators dataset.

Health promotion and education (HPE) needs to be evaluated on a national scale. This note draws attention to the existence, possible uses and pitfalls of a little known dataset which provides information on English district health authorities' HPE expenditure for the first time. Despite its problems, cautious uses of this data has the potential to significantly increase the knowledge and understanding of local level HPE in England.

Databases, Factual↗

Evaluating the impact of the National Healthy School Standard: using national datasets.

An evaluation of the National Healthy School Standard (NHSS) was undertaken by the authors on behalf of the Department of Health and the Department for Education and Skills. One part of the evaluation involved gaining access to a number of datasets derived from previous research and analysing the health-related outcomes of schools which had attained Level 3 of the NHSS, compared with those of other schools. The sources which provided the most interesting findings were the Health-Related Behaviour Questionnaire (HRBQ) survey and the Ofsted database of school inspection ratings. This paper describes the statistical methods used, and the results of the HRBQ and Ofsted analyses. Using HRBQ data, many pupil-level outcomes were explored, but relatively few indicated significant differences and even those tended to be quite small. The Ofsted school-level data yielded stronger evidence of NHSS impact. The paper concludes by suggesting possible reasons for these findings.

Adolescent↗

Feature expressions: creating and manipulating sequence datasets.

Annotation of features, such as introns, exons and protein coding regions in GenBank/EMBL/DDBJ entries is now standardized through use of the Features Table (FT) language. The essence of the FT language is described by the relation 'expression-->sequence', meaning that each FT expression evaluates to a sequence. For example, the expression M74750:1..50 evaluates to the first 50 bases of the sequence with accession number M74750. Because FT is intrinsic to the database definition, it can serve as a software- and platform-independent lingua franca for sequence manipulation. The XYLEM package makes it possible to create and manipulate sequence datasets using FT expressions. FEATURES is a program that resolves FT expressions into their corresponding sequences. Annotated features can be retrieved either by feature key or by expression. Even unannotated portions of a sequence can be retrieved by user-generated FT expressions. Applications of the FT language include retrieval of subsequences from large sequence entries, generation of chromosome models or artificial DNA constructs, and representation of restriction maps or mutants.

Base Sequence↗

The mitBASE human dataset structure.

MitBASE is a comprehensive and integrated mitochondrial genome database funded within the EU BIOTECH PROGRAM. It is a project for the development and implementation of an integrated and comprehensive database of mitochondrial data which will collect all available information from different organisms and from intraspecies variants and mutants. The present paper describes the structure of the Human dataset in mitBASE where human molecular data are distinguished from clinical and pathological data. MitBASE home page address is: http://www.ebi.ac.uk/htbin/Mitbase/mitb ase.pl

Computer Communication Networks↗

A new approach for filtering noise from high-density oligonucleotide microarray datasets.

Although DNA microarrays are powerful tools for profiling gene expression, the dynamic range and the sheer number of signals produced require efficient procedures for distinguishing false positive results (noise) from changes in expression that are 'real' (independently reproducible). We have developed an approach to filter noise from datasets generated when high density oligonucleotide-based microarrays are used to compare two distinct RNA populations. First, we performed comparisons between chips hybridized with cRNAs prepared from an identical starting RNA population; an 'Increase' or 'Decrease' call in such a comparison was defined as a false positive. Plotting the average distribution of these false positive signal intensities across 18 such comparisons of nine independent RNA preparations allowed us to develop a series of noise-filtering look-up tables (LUTs). Using a database of 70 separate chip-to-chip comparisons between distinct RNA preparations prepared by different workers at different sites and at different times, we show that the LUTs can be used to predict the likelihood that a given transcript called Increased or Decreased in one comparison will again be called Increased or Decreased in a replicate comparison. Evidence is presented that this LUT-based scoring system provides greater predictive value for reproducible microarray results than imposition of arbitrary fold-change thresholds and accurately predicts which microarray-identified changes will be validated by independent assays such as quantitative real-time PCR.

Databases as Topic↗

PDA v.2: improving the exploration and estimation of nucleotide polymorphism in large datasets of heterogeneous DNA.

Pipeline Diversity Analysis (PDA) is an open-source, web-based tool that allows the exploration of polymorphism in large datasets of heterogeneous DNA sequences, and can be used to create secondary polymorphism databases for different taxonomic groups, such as the Drosophila Polymorphism Database (DPDB). A new version of the pipeline presented here, PDA v.2, incorporates substantial improvements, including new methods for data mining and grouping sequences, new criteria for data quality assessment and a better user interface. PDA is a powerful tool to obtain and synthesize existing empirical evidence on genetic diversity in any species or species group. PDA v.2 is available on the web at http://pda.uab.es/.

Algorithms↗

Evaluating Language Models for Biomedical Fact-Checking: A Benchmark Dataset for Cancer Variant Interpretation Verification.

Accurate interpretation of genomic variants is critical for precision oncology but remains slow and dependent on specialized expertise. Public knowledgebases such as the Clinical Interpretation of Variants in Cancer (CIViC) help by curating literature-backed variant interpretations in a structured form, yet verification and review have become major bottlenecks. To address this, we developed CIViC-Fact, a benchmark dataset and pipeline for testing automated systems that verify the accuracy of cancer variant claims. CIViC-Fact links structured claims to sentence-level supporting or refuting evidence from full-text articles, and includes expert annotations and explanations. We evaluated multiple language models. Proprietary models performed well without training, but a smaller open-source model, fine-tuned on CIViC-Fact, achieved the highest accuracy (89%). Applying our fact-checking pipeline to real CIViC entries showed that reviewing less than 20% of content, focusing on flagged entries, would be sufficient to catch over half of all errors. This AI-assisted triage greatly accelerates the review process without replacing or reducing expert insight, ensuring that existing careful oversight remains in place while curators can work more efficiently. CIViC-Fact provides a realistic, high-consequence framework for biomedical fact-checking and a path toward more rigorous and efficient knowledgebase curation.

Journal Article↗

Range-dependent random graphs and their application to modeling large small-world Proteome datasets.

In this paper we consider the problem of characterizing and modeling large-scale networks using classes of range-dependent graphs which possess appropriate small-world properties. The application we have in mind is to bioinformatics, where methods of rapid protein identification mean that such proteome datasets, listing various observed protein-protein associations, will become more and more prevalent. We introduce a class of range-dependent graphs, governed by a power law relating intervertex range to edge probability, which are amenable to analysis, and for which macroscopic graph parameters are given by explicit forms. We show how these may be employed in representing a given network using a maximum likelihood approach. This in turn annotates every given edge with its range, representing the tendency for such an association to be transitive. We apply this technique to published proteome data, and demonstrate that known protein associations are thus identified.

Journal Article↗

Wavelet coding of volumetric medical datasets.

Several techniques based on the three-dimensional (3-D) discrete cosine transform (DCT) have been proposed for volumetric data coding. These techniques fail to provide lossless coding coupled with quality and resolution scalability, which is a significant drawback for medical applications. This paper gives an overview of several state-of-the-art 3-D wavelet coders that do meet these requirements and proposes new compression methods exploiting the quadtree and block-based coding concepts, layered zero-coding principles, and context-based arithmetic coding. Additionally, a new 3-D DCT-based coding scheme is designed and used for benchmarking. The proposed wavelet-based coding algorithms produce embedded data streams that can be decoded up to the lossless level and support the desired set of functionality constraints. Moreover, objective and subjective quality evaluation on various medical volumetric datasets shows that the proposed algorithms provide competitive lossy and lossless compression results when compared with the state-of-the-art.

Algorithms↗

The Chinese Visible Human (CVH) datasets incorporate technical and imaging advances on earlier digital humans.

We report the availability of a digitized Chinese male and a digitzed Chinese female typical of the population and with no obvious abnormalities. The embalming and milling procedures incorporate three technical improvements over earlier digitized cadavers. Vascular perfusion with coloured gelatin was performed to facilitate blood vessel identification. Embalmed cadavers were embedded in gelatin and cryosectioned whole so as to avoid section loss resulting from cutting the body into smaller pieces. Milling performed at -25 degrees C prevented small structures (e.g. teeth, concha nasalis and articular cartilage) from falling off from the milling surface. The male image set (.tiff images each of 36 Mb) has a section resolution of 3072 x 2048 pixels ( approximately 170 micro m, the accompanying magnetic resonance imaging and computer tomography data have a resolution of 512 x 512, i.e. approximately 440 micro m). The Chinese Visible Human male and female datasets are available at http://www.chinesevisiblehuman.com. (The male is 90.65 Gb and female 131.04 Gb). MPEG videos of direct records of real-time volume rendering are at: http://www.cse.cuhk.edu.hk/~crc

Adult↗

Mucous membrane and lower respiratory building related symptoms in relation to indoor carbon dioxide concentrations in the 100-building BASE dataset.

UNLABELLED: Indoor air pollutants are a potential cause of building related symptoms and can be reduced by increasing ventilation rates. Indoor carbon dioxide (CO(2)) concentration is an approximate surrogate for concentrations of occupant-generated pollutants and for ventilation rate per occupant. Using the US EPA 100 office-building BASE Study dataset, we conducted multivariate logistic regression analyses to quantify the relationship between indoor CO(2) concentrations (dCO(2)) and mucous membrane (MM) and lower respiratory system (LResp) building related symptoms, adjusting for age, sex, smoking status, presence of carpet in workspace, thermal exposure, relative humidity, and a marker for entrained automobile exhaust. In addition, we tested the hypothesis that certain environmentally mediated health conditions (e.g., allergies and asthma) confer increased susceptibility to building related symptoms. Adjusted odds ratios (ORs) for statistically significant, dose-dependent associations (P < 0.05) for combined mucous membrane, dry eyes, sore throat, nose/sinus congestion, sneeze, and wheeze symptoms with 100 p.p.m. increases in dCO(2) ranged from 1.1 to 1.2. Building occupants with certain environmentally mediated health conditions were more likely to report that they experience building related symptoms than those without these conditions (statistically significant ORs ranged from 1.5 to 11.1, P < 0.05). PRACTICAL IMPLICATIONS: These results suggest that provision of sufficient per-person outdoor ventilation air, could significantly decrease prevalence of selected building related symptoms. The observed relationship between indoor minus outdoor CO(2) concentrations and mucous membrane and lower respiratory symptoms suggests that air contaminants are implicated in the etiology of building related symptoms. Levels of indoor air pollutants that are suspected to cause building related symptoms could be reduced by increasing ventilation rates, improving ventilation effectiveness, or reducing sources of indoor air pollutants, if known.

Adult↗