Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Feature selection with limited datasets.

Computer-aided diagnosis has the potential of increasing diagnostic accuracy by providing a second reading to radiologists. In many computerized schemes, numerous features can be extracted to describe suspect image regions. A subset of these features is then employed in a data classifier to determine whether the suspect region is abnormal or normal. Different subsets of features will, in general, result in different classification performances. A feature selection method is often used to determine an "optimal" subset of features to use with a particular classifier. A classifier performance measure (such as the area under the receiver operating characteristic curve) must be incorporated into this feature selection process. With limited datasets, however, there is a distribution in the classifier performance measure for a given classifier and subset of features. In this paper, we investigate the variation in the selected subset of "optimal" features as compared with the true optimal subset of features caused by this distribution of classifier performance. We consider examples in which the probability that the optimal subset of features is selected can be analytically computed. We show the dependence of this probability on the dataset sample size, the total number of features from which to select, the number of features selected, and the performance of the true optimal subset. Once a subset of features has been selected, the parameters of the data classifier must be determined. We show that, with limited datasets and/or a large number of features from which to choose, bias is introduced if the classifier parameters are determined using the same data that were employed to select the "optimal" subset of features.

Bias↗

OligoSpawn: a software tool for the design of overgo probes from large unigene datasets.

BACKGROUND: Expressed sequence tag (EST) datasets represent perhaps the largest collection of genetic information. ESTs can be exploited in a variety of biological experiments and analysis. Here we are interested in the design of overlapping oligonucleotide (overgo) probes from large unigene (EST-contigs) datasets. RESULTS: OLIGOSPAWN is a suite of software tools that offers two complementary services, namely (1) the selection of "unique" oligos each of which appears in one unigene but does not occur (exactly or approximately) in any other and (2) the selection of "popular" oligos each of which occurs (exactly or approximately) in as many unigenes as possible. In this paper, we describe the functionalities of OLIGOSPAWN and the computational methods it employs, and we report on experimental results for the overgo probes designed with it. CONCLUSION: The algorithms we designed are highly efficient and capable of processing unigene datasets of sizes on the order of several tens of Mb in a few hours on a regular PC. The software has been used to design overgo probes employed to screen a barley BAC library (Hordeum vulgare). OLIGOSPAWN is freely available at http://oligospawn.ucr.edu/.

Base Sequence↗

Preferred analysis methods for Affymetrix GeneChips revealed by a wholly defined control dataset.

BACKGROUND: As more methods are developed to analyze RNA-profiling data, assessing their performance using control datasets becomes increasingly important. RESULTS: We present a 'spike-in' experiment for Affymetrix GeneChips that provides a defined dataset of 3,860 RNA species, which we use to evaluate analysis options for identifying differentially expressed genes. The experimental design incorporates two novel features. First, to obtain accurate estimates of false-positive and false-negative rates, 100-200 RNAs are spiked in at each fold-change level of interest, ranging from 1.2 to 4-fold. Second, instead of using an uncharacterized background RNA sample, a set of 2,551 RNA species is used as the constant (1x) set, allowing us to know whether any given probe set is truly present or absent. Application of a large number of analysis methods to this dataset reveals clear variation in their ability to identify differentially expressed genes. False-negative and false-positive rates are minimized when the following options are chosen: subtracting nonspecific signal from the PM probe intensities; performing an intensity-dependent normalization at the probe set level; and incorporating a signal intensity-dependent standard deviation in the test statistic. CONCLUSIONS: A best-route combination of analysis methods is presented that allows detection of approximately 70% of true positives before reaching a 10% false-discovery rate. We highlight areas in need of improvement, including better estimate of false-discovery rates and decreased false-negative rates.

Algorithms↗

Identification of T cell-restricted genes, and signatures for different T cell responses, using a comprehensive collection of microarray datasets.

We used a comprehensive collection of Affymetrix microarray datasets to ascertain which genes or molecules distinguish the known major subsets of human T cells. Our strategy allowed us to identify the genes expressed in most T cell subsets: TCR alphabeta+ and gammadelta+, three effector subsets (Th1, Th2, and T follicular helper cells), T central memory, T effector memory, activated T cells, and others. Our genechip dataset also allowed for identification of genes preferentially or exclusively expressed by T cells, compared with numerous non-T cell leukocyte subsets profiled. Cross-comparisons between microarray datasets revealed important features of certain subsets. For instance, blood gammadelta T cells expressed no unique gene transcripts, but did differ from alphabeta T cells in numerous genes that were down-regulated. Hierarchical clustering of all the genes differentially expressed between T cell subsets enabled the identification of precise signatures. Moreover, the different T cell subsets could be distinguished at the level of gene expression by a smaller subset of predictor genes, most of which have not previously been associated directly with any of the individual subsets. T cell activation had the greatest influence on gene regulation, whereas central and effector memory T cells displayed surprisingly similar gene expression profiles. Knowledge of the patterns of gene expression that underlie fundamental T cell activities, such as activation, various effector functions, and immunological memory, provide the basis for a better understanding of T cells and their role in immune defense.

Cluster Analysis↗

Construction of dataset for Virtual Chinese Male No.1.

OBJECTIVE: To establish digitized Virtual Chinese Human Male No.1 (VCH-M1) image dataset with a 0.2-mm equal interval. METHODS: The body of a 24-year-old male was used for this study. Perfusion with phenol and vermilion of the arteries was performed, followed by body shape adjustment by cold saline and pre-embedding with broken ices in an upside-down position, which was completed in a stepwise procedure to minimize body shape deformation. Section milling was conducted subsequently with the section thickness of 2 mm and the section images were captured by digital camera, which were immediately transferred to a computer for storage and processing. RESULTS: A total of 9 232 sections were obtained for the whole body, and the resolution of each of the image in TIF format was 3 024x2 016 pixels, resulting in the size of approximately 18 M for each image and about 161 G for the whole dataset. CONCLUSIONS: Compared with VCH-F1, the image quality in VCH-M1 dataset is significantly improved, demonstrated by much clearer tissue boundary in the images and minimized body shape deformation during the embedding process.

Adult↗

[Establishment of Internet-based database of the Virtual Chinese Human dataset].

To establish an Internet-based database for the dataset of Virtual Chinese Human that is accessible to the interested researchers, modifications and compression of the original VCH-format dataset of Virtual Chinese Human were performed before it was uploaded to the server, and RAID0+1 storage technology was adopted with specific download accesses designed for different users. After dataset modification and compression, the data size was considerably reduced to allow convenient data storage and transfer. The RAID0+1 storage technology guarantees the security and high-speed download of data through different means established. Internet-based database provides important accesses for sharing the achievement in virtual human study between world-wide researchers, which has been imperative in the present situation of science development.

Anatomy, Cross-Sectional↗

Cluster analysis of Wisconsin Breast Cancer dataset using self-organizing maps.

This work deals with multidimensional data analysis, precisely cluster analysis applied to a very well known dataset, the Wisconsin Breast Cancer dataset. After the introduction of the topics of the paper the cluster analysis concept is shortly explained and different methods of cluster analysis are compared. Further, the Kohonen model of self-organizing maps is briefly described together with an example and with explanations of how the cluster analysis can be performed using the maps. After describing the data set and the methodology used for the analysis we present the findings using textual as well as visual descriptions and conclude that the approach is a useful complement for assessing multidimensional data and that this dataset has been overused for automated decision benchmarking purposes, without a thorough analysis of the data it contains.

Breast Neoplasms↗

Expression regulation network in papillae of sea cucumbers: Whole-transcriptome and DNA methylation datasets.

To elucidate the expression regulation network of papilla size of sea cucumbers (Apostichopus japonicus), the whole-transcriptome and DNA methylome datasets of different sizes of papillae in sea cucumbers were generated. Average clean bases of whole-transcriptome (16.35 G) and DNA methylome (28.92 G) were obtained using RNA sequencing and whole-genome bisulfite sequencing techniques. A total of 3,188 ceRNA networks were also identified including 3,081 long non-coding RNAs (lncRNA)/microRNAs (miRNA)/mRNA networks and 107 circular RNA (circRNA)/miRNA/mRNA networks. Methylome data indicate that there were 3,307 and 3,776 differentially methylated regions (DMRs) with high-level methylation as well as 3,125 and 3,016 DMRs with low-level methylation in big papillae compared to small papillae. The identified DMRs were mainly distributed in introns, promotors, or exons. The whole-transcriptome and DNA methylome datasets generated from this study not only established a robust theoretical foundation (especially from the epigenetic aspect) for elucidating expression regulation network determining papilla size in sea cucumbers but also can be a valuable resource of biomarker mining for papilla appearance-based selective breeding in sea cucumbers.

DNA Methylation↗

Identification of discriminators of hepatoma by gene expression profiling using a minimal dataset approach.

The severity of hepatocellular carcinoma (HCC) and the lack of good diagnostic markers and treatment strategies have rendered the disease a major challenge. Previous microarray analyses of HCC were restricted to the selected tissue sample sets without validation on an independent series of tissue samples. We describe an approach to the identification of a composite discriminator cassette by intersecting different microarray datasets. We studied the global transcriptional profiles of matched HCC tumor and nontumor liver samples from 37 patients using cDNA (cDNA) microarrays. Application of nonparametric Wilcoxon statistical analyses (P < 1 x 10(-6)) and the criteria of 1.5-fold differential gene expression change resulted in the identification of 218 genes, including BMI-1, ERBB3, and those involved in the ubiquitin-proteasome pathway. Elevated ERBB2 and epidermal growth factor receptor (EGFR) expression levels were detected in ERBB3-expressing tumors, suggesting the presence of ERBB3 cognate partners. Comparison of our dataset with an earlier study of approximately 150 tissue sets identified multiple overlapping discriminator markers, suggesting good concordance of data despite differences in patient populations and technology platforms. These overlapping discriminator markers could distinguish HCC tumor from nontumor liver samples with reasonable precision and the features were unlikely to appear by chance, as measured by Monte Carlo simulations. More significantly, validation of the discriminator cassettes on an independent set of 58 liver biopsy specimens yielded greater than 93% prediction accuracy. In conclusion, these data indicate the robustness of expression profiling in marker discovery using limited patient tissue specimens as well as identify novel genes that are highly likely to be excellent markers for HCC diagnosis and treatment.

Biopsy↗

The interobserver reliability and validity of volume calculation from three-dimensional ultrasound datasets in the in vitro setting.

OBJECTIVES: The primary aim of this validation study was to determine the interobserver reliability and validity of measurements of phantom objects of known volume using conventional and rotational techniques of volume calculation according to measurement technique. METHODS: Two observers each acquired a single three-dimensional ultrasound dataset of three water-filled objects of different size and shape. The same two observers measured all six datasets using both the conventional technique and the newer rotational technique (Virtual Organ Computer-aided AnaLysis, VOCAL( trade mark )) of volume calculation. Reliability was assessed by calculating intraclass correlation coefficients (ICC) and validity by examining the percentage difference from the 'true' volume, as determined by a water displacement technique, by the limits of agreement method. RESULTS: All of the techniques were highly reliable (ICC: 0.9962-0.9997) and valid to within 4% of the 'true' volumes. There were no significant differences in reliability according to measurement plane or between observers. Measurements made with the 6 degrees rotation step were significantly more reliable than those made by all other techniques with the exception of the 9 degrees rotation step (P < 0.05) and significantly more valid than those made with the 30 degrees rotation step or conventional technique (P < 0.05). CONCLUSIONS: Volume calculation in the in vitro setting is both reliable and valid but is dependent upon the technique applied, with rotational measurements of volume proving superior to conventional techniques.

Female↗

Phylogenetic analysis of a dataset of fungal 5.8S rDNA sequences shows that highly divergent copies of internal transcribed spacers reported from Scutellospora castanea are of ascomycete origin.

Using a dataset comprising 5.8S rDNA sequences from a wide range of fungi, we show that some sequences reported recently from the arbuscular mycorrhizal (AM) fungus Scutellospora castanea most likely originate from Ascomycetes. Other ITS and 5.8S sequences which were previously reported are confirmed as being clearly of mycorrhizal origin and are variable within one isolate of S. castanea. However, these results mean that previous conclusions which were drawn regarding the heterokaryotic status of AM fungal spores remain unproven. We provide an enlarged 5.8S rDNA dataset that can be used to check ITS sequences for conflicts with well-established phylogenies of the organisms that they were obtained from.

Amino Acid Sequence↗

The centroidal algorithm in molecular similarity and diversity calculations on confidential datasets.

Chemical structure provides exhaustive description of a compound, but it is often proprietary and thus an impediment in the exchange of information. For example, structure disclosure is often needed for the selection of most similar or dissimilar compounds. Authors propose a centroidal algorithm based on structural fragments (screens) that can be efficiently used for the similarity and diversity selections without disclosing structures from the reference set. For an increased security purposes, authors recommend that such set contains at least some tens of structures. Analysis of reverse engineering feasibility showed that the problem difficulty grows with decrease of the screen's radius. The algorithm is illustrated with concrete calculations on known steroidal, quinoline, and quinazoline drugs. We also investigate a problem of scaffold identification in combinatorial library dataset. The results show that relatively small screens of radius equal to 2 bond lengths perform well in the similarity sorting, while radius 4 screens yield better results in diversity sorting. The software implementation of the algorithm taking SDF file with a reference set generates screens of various radii which are subsequently used for the similarity and diversity sorting of external SDFs. Since the reverse engineering of the reference set molecules from their screens has the same difficulty as the RSA asymmetric encryption algorithm, generated screens can be stored openly without further encryption. This approach ensures an end user transfers only a set of structural fragments and no other data. Like other algorithms of encryption, the centroid algorithm cannot give 100% guarantee of protecting a chemical structure from dataset, but probability of initial structure identification is very small-order of 10(-40) in typical cases.

Algorithms↗

Comprehensive comparisons of the current human, mouse, and rat RefSeq, Ensembl, EST, and FANTOM3 datasets: identification of new human genes with specific tissue expression profile.

Our understanding of functional genetic elements in the genomes is continuously growing and new entries are entered in various databases on a regular basis. We have here merged the genetic elements in RefSeq, Ensembl, FANTOM3, HINV, and NCBI:s ESTdb using the genome assemblies in order to achieve a comprehensive picture of the current status of the identity and gene number in human, mouse, and rat. The number of human protein coding genes has not increased (25,043) while the increased sequencing of mouse transcripts has provided the considerably higher number of protein coding genes (31,578) in mouse. The results indicate large discrepancies between the datasets, as considerable numbers of unique transcripts can be found in each dataset. Despite the high number of ncRNA (38,129 in mouse) there are also almost 20,000 EST clusters in both mouse and humans with more than one EST that do not overlap any transcript suggesting that several new genetic elements are still to be found. We also demonstrated presence of new genes by identifying new human ones that have specific tissue profiles, using RT-PCR on rat tissues.

Animals↗

A dataset of estimated heterozygous individual and carrier couple frequencies for pan-ancestry carrier screening.

The data described in this publication supported the development and evaluation of pan-ancestry reproductive carrier screening panels for autosomal recessive (AR) and X-linked (XL) conditions. Raw data included combined sets of DNA variants in 1,350 AR/XL genes obtained from the ClinVar and gnomAD databases. The dataset enabled calculations of positive yield for individuals and couples across both ancestry-specific and pan-ancestry, optimised "Goldilocks"-ranked gene panels, addressing population-specific variations in the frequencies of heterozygous individuals and carrier couples. The positive yield analysis offered a performance metric for carrier screening panels, facilitating the modeling of screening performance for panels of varying sizes and composition and providing resources for optimizing panel content to ensure equity across underrepresented genetic ancestries The dataset can support ongoing research into the equitable application of carrier screening and offers significant reuse potential for refining population genetic screening practices, validating computational models, and developing frameworks to update carrier screening panels in alignment with evolving genomic data, including in underrepresented and minority populations.

Carrier screening↗

Period analysis of cancer patient survival in datasets from which the month of diagnosis has been removed.

Up-to-date monitoring of long-term survival is an important task of population-based cancer registries. Period analysis, a new method of survival analysis introduced a few years ago, has been shown to be particularly useful for that purpose. The "classical" period analysis uses a life-table approach which requires both the year and month of diagnosis for implementation in pertinent software programs. However, an increasing number of cancer registries remove the month of diagnosis from their datasets, mainly to ensure the highest possible protection against re-identification of patients. In this paper, we present modifications of period analysis that allow the application of this technique, while almost completely preserving its advantages, in datasets without the month of diagnosis. The modified techniques are illustrated and evaluated using examples from the Surveillance, Epidemiology, and End Results (SEER) programme of the United States (US) National Cancer Institute (NCI), which also has removed month of diagnosis from its most recently released public use database.

Adolescent↗

An integrated taxonomic study of Fusarium langsethiae, Fusarium poae and Fusarium sporotrichioides based on the use of composite datasets.

An integrated systematic study was carried out to clarify the taxonomical position and relationship of Fusarium langsethiae to other taxa within the Fusarium section Sporotrichiella. Strains of this species were compared with strains of the closely related species Fusarium poae and Fusarium sporotrichioides using a composite dataset. This set consisted of DNA sequences derived from the ribosomal internal transcribed spacer (ITS) regions, partial sequences of the ribosomal intergenic spacer (IGS) region, the beta-tubulin and translation elongation factor-1 alpha (EF-1alpha) genes, AFLP fingerprints, chromatographic data on secondary metabolites and morphological data and growth characteristics. From these combined data, a consensus matrix was calculated by taking the mean of all pairwise distances between single isolates over all separate datasets. The consensus matrix was used as the basis for the construction of a UPGMA dendrogram and a multidimensional scaling, both of which revealed a clear separation of the three taxa. Partial IGS, EF-1alpha and beta-tubulin sequence-as well as chromatography-and AFLP-derived similarities turned out to be comparably consistent, while ITS sequence- and morphology-derived similarity matrices were rather divergent.

Base Sequence↗

Dynamics of EF-G interaction with the ribosome explored by classification of a heterogeneous cryo-EM dataset.

A method of supervised classification using two available structure templates was applied to investigate the possible heterogeneity existing in a large cryo-EM dataset of an Escherichia coli 70S ribosome-EF-G complex. Two subpopulations showing the ribosome in distinct conformational states, related by a ratchet-like rotation of the 30S subunit with respect to the 50S subunit, were extracted from the original dataset. The possible presence of additional intermediate states is discussed.

Cryoelectron Microscopy↗

Individual Differences in Cognitive Aging Rodent Datasets (ID-CARD): A collaborative platform for behavioral analysis across the lifespan.

Understanding cognitive aging requires approaches that capture individual variability while enabling integration across studies. In rodent models, behavioral data are central to this effort, yet cross-laboratory differences in experimental design limit comparability and constrain secondary analysis. To address this gap, we developed the Individual Differences in Cognitive Aging Rodent Datasets (ID-CARD), a first-of-its-kind collaborative repository aggregating trial-level Morris water maze data from multiple laboratories. ID-CARD is designed to support large-scale, integrative analyses and to facilitate secondary use of existing behavioral data in alignment with emerging data-sharing and transparency initiatives. Rather than imposing retrospective harmonization of experimental protocols, we implemented a normalization and modeling framework that enables comparison of learning trajectories while preserving meaningful variation across studies. Behavioral data from >&#x202f;5000 rats spanning common strains, both sexes, and multiple ages were normalized in training and performance domains and fit with a logarithmic function to derive an error accumulation rate coefficient (EARC) as a measure of spatial learning. Age was strongly associated with increased EARC, indicating attenuated learning, even after adjusting for non-spatial cue performance. Analyses of goodness of fit revealed systematic structure in learning dynamics, where age was associated with reduced learning-curve conformity after accounting for overall performance. Inter-individual variability in spatial learning also increased with age, with strain-specific interactions. These findings demonstrate that integrated analysis of heterogeneous behavioral datasets can yield robust, individual-level insights into cognitive aging. ID-CARD provides a scalable resource and analytic framework to advance discovery in behavioral neuroscience by enabling reuse, integration, and comparative analysis of existing data.

Cognitive aging↗