Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Acquiring a four-dimensional computed tomography dataset using an external respiratory signal.

Four-dimensional (4D) methods strive to achieve highly conformal radiotherapy, particularly for lung and breast tumours, in the presence of respiratory-induced motion of tumours and normal tissues. Four-dimensional radiotherapy accounts for respiratory motion during imaging, planning and radiation delivery, and requires a 4D CT image in which the internal anatomy motion as a function of the respiratory cycle can be quantified. The aims of our research were (a) to develop a method to acquire 4D CT images from a spiral CT scan using an external respiratory signal and (b) to examine the potential utility of 4D CT imaging. A commercially available respiratory motion monitoring system provided an 'external' tracking signal of the patient's breathing. Simultaneous recording of a TTL 'X-Ray ON' signal from the CT scanner indicated the start time of CT image acquisition, thus facilitating time stamping of all subsequent images. An over-sampled spiral CT scan was acquired using a pitch of 0.5 and scanner rotation time of 1.5 s. Each image from such a scan was sorted into an image bin that corresponded with the phase of the respiratory cycle in which the image was acquired. The complete set of such image bins accumulated over a respiratory cycle constitutes a 4D CT dataset. Four-dimensional CT datasets of a mechanical oscillator phantom and a patient undergoing lung radiotherapy were acquired. Motion artefacts were significantly reduced in the images in the 4D CT dataset compared to the three-dimensional (3D) images, for which respiratory motion was not accounted. Accounting for respiratory motion using 4D CT imaging is feasible and yields images with less distortion than 3D images. 4D images also contain respiratory motion information not available in a 3D CT image.

Algorithms↗

Estimating dataset size requirements for classifying DNA microarray data.

A statistical methodology for estimating dataset size requirements for classifying microarray data using learning curves is introduced. The goal is to use existing classification results to estimate dataset size requirements for future classification experiments and to evaluate the gain in accuracy and significance of classifiers built with additional data. The method is based on fitting inverse power-law models to construct empirical learning curves. It also includes a permutation test procedure to assess the statistical significance of classification performance for a given dataset size. This procedure is applied to several molecular classification problems representing a broad spectrum of levels of complexity.

Algorithms↗

De novo clustering of large long-read transcriptome datasets with isONclust3.

MOTIVATION: Long-read sequencing techniques can sequence transcripts from end to end, greatly improving our ability to study the transcription process. Although there are several well-established tools for long-read transcriptome analysis, most are reference-based. This limits the analysis of organisms without high-quality reference genomes and samples or genes with high variability (e.g. cancer samples or some gene families). In such settings, analysis using a reference-free method is favorable. The computational problem of clustering long reads by region of common origin is well-established for reference-free transcriptome analysis pipelines. Such clustering enables large datasets to be split roughly by gene family and, therefore, an independent analysis of each cluster. There exist tools for this. However, none of those tools can efficiently process the large amount of reads that are now generated by long-read sequencing technologies. RESULTS: We present isONclust3, an improved algorithm over isONclust and isONclust2, to cluster massive long-read transcriptome datasets into gene families. Like isONclust, isONclust3 represents each cluster with a set of minimizers. However, unlike other approaches, isONclust3 dynamically updates the cluster representation during clustering by adding high-confidence minimizers from new reads assigned to the cluster and employs an iterative cluster-merging step. We show that isONclust3 yields results with higher or comparable quality to state-of-the-art algorithms but is 10-100 times faster on large datasets. Also, using a 256 Gb computing node, isONclust3 was the only tool that could cluster 37 million PacBio reads, which is a typical throughput of the recent PacBio Revio sequencing machine. AVAILABILITY AND IMPLEMENTATION: https://github.com/aljpetri/isONclust3.

Algorithms↗

Combining Annotation Software to Identify Orthologous Genes (CASIO) Provides a New Dataset of Orthologous Genes for Swallowtail Butterflies.

With the massive increase in genomic resources, it is becoming increasingly popular to analyse thousands of loci across many species. However, many of the available genomes are not annotated, which hinders an efficient search for orthologous protein-coding genes. Here, we aim to develop a semi-automated pipeline and compare four genomic annotation methods (BRAKER2, BUSCO, Miniprot and Scipio). Our results highlight the importance of integrating multiple annotation tools to optimise ortholog detection and improve genomic studies. Each annotation method showed different strengths. BRAKER2 annotated a substantial number of genes. BUSCO, despite limitations inherent to its reference database, identified a higher number of orthologs. Miniprot exhibited notable flexibility in accommodating diverse protein datasets, whereas Scipio successfully recovered a considerable set of genes that were not detected by the other tools. The combination of these tools allowed for more comprehensive ortholog detection. Taking advantage of this pipeline, we developed a comprehensive dataset of orthologous genes for swallowtail butterflies (Lepidoptera: Papilionidae), called Papilionidae_odb, which will facilitate future studies, especially for a non-model group with abundant genomic data and few transcriptomic resources. We tested Papilionidae_odb by inferring a robust phylogenetic framework for Leptocircini using 142 complete genomes, which improved branch support for some phylogenetic relationships, although challenges remained in resolving relationships within certain species groups, likely due to rapid radiations. Our results highlight the complementary nature of the annotation methods and suggest that combining these tools can yield more accurate results in genomic research. This approach was implemented in a Snakemake workflow called CASIO (Combining Annotation Software to Identify Orthologous genes) and can easily be applied to other non-model groups to improve genomic datasets in diverse taxa where transcriptomic resources are still limited.

Animals↗

Combining datasets to predict the effects of regulation of environmental lead exposure in housing stock.

A model for children's blood lead concentrations as a function of environmental lead exposures was developed by combining two nationally representative sources of data that characterize the marginal distributions of blood lead and environmental lead with a third regional dataset that contains joint measures of blood lead and environmental lead. The complicating factor addressed in this article was the fact that methods for assessing environmental lead were different in the national and regional datasets. Relying on an assumption of transportability (that although the marginal distributions of blood lead and environmental lead may be different between the regional dataset and the nation as a whole, the joint relationship between blood lead and environmental lead is the same), the model makes use of a latent variable approach to estimate the joint distribution of blood lead and environmental lead nationwide.

Biometry↗

Feature selection with limited datasets.

Computer-aided diagnosis has the potential of increasing diagnostic accuracy by providing a second reading to radiologists. In many computerized schemes, numerous features can be extracted to describe suspect image regions. A subset of these features is then employed in a data classifier to determine whether the suspect region is abnormal or normal. Different subsets of features will, in general, result in different classification performances. A feature selection method is often used to determine an "optimal" subset of features to use with a particular classifier. A classifier performance measure (such as the area under the receiver operating characteristic curve) must be incorporated into this feature selection process. With limited datasets, however, there is a distribution in the classifier performance measure for a given classifier and subset of features. In this paper, we investigate the variation in the selected subset of "optimal" features as compared with the true optimal subset of features caused by this distribution of classifier performance. We consider examples in which the probability that the optimal subset of features is selected can be analytically computed. We show the dependence of this probability on the dataset sample size, the total number of features from which to select, the number of features selected, and the performance of the true optimal subset. Once a subset of features has been selected, the parameters of the data classifier must be determined. We show that, with limited datasets and/or a large number of features from which to choose, bias is introduced if the classifier parameters are determined using the same data that were employed to select the "optimal" subset of features.

Bias↗

Construction of dataset for Virtual Chinese Male No.1.

OBJECTIVE: To establish digitized Virtual Chinese Human Male No.1 (VCH-M1) image dataset with a 0.2-mm equal interval. METHODS: The body of a 24-year-old male was used for this study. Perfusion with phenol and vermilion of the arteries was performed, followed by body shape adjustment by cold saline and pre-embedding with broken ices in an upside-down position, which was completed in a stepwise procedure to minimize body shape deformation. Section milling was conducted subsequently with the section thickness of 2 mm and the section images were captured by digital camera, which were immediately transferred to a computer for storage and processing. RESULTS: A total of 9 232 sections were obtained for the whole body, and the resolution of each of the image in TIF format was 3 024x2 016 pixels, resulting in the size of approximately 18 M for each image and about 161 G for the whole dataset. CONCLUSIONS: Compared with VCH-F1, the image quality in VCH-M1 dataset is significantly improved, demonstrated by much clearer tissue boundary in the images and minimized body shape deformation during the embedding process.

Adult↗

[Establishment of Internet-based database of the Virtual Chinese Human dataset].

To establish an Internet-based database for the dataset of Virtual Chinese Human that is accessible to the interested researchers, modifications and compression of the original VCH-format dataset of Virtual Chinese Human were performed before it was uploaded to the server, and RAID0+1 storage technology was adopted with specific download accesses designed for different users. After dataset modification and compression, the data size was considerably reduced to allow convenient data storage and transfer. The RAID0+1 storage technology guarantees the security and high-speed download of data through different means established. Internet-based database provides important accesses for sharing the achievement in virtual human study between world-wide researchers, which has been imperative in the present situation of science development.

Anatomy, Cross-Sectional↗

Expression regulation network in papillae of sea cucumbers: Whole-transcriptome and DNA methylation datasets.

To elucidate the expression regulation network of papilla size of sea cucumbers (Apostichopus japonicus), the whole-transcriptome and DNA methylome datasets of different sizes of papillae in sea cucumbers were generated. Average clean bases of whole-transcriptome (16.35 G) and DNA methylome (28.92 G) were obtained using RNA sequencing and whole-genome bisulfite sequencing techniques. A total of 3,188 ceRNA networks were also identified including 3,081 long non-coding RNAs (lncRNA)/microRNAs (miRNA)/mRNA networks and 107 circular RNA (circRNA)/miRNA/mRNA networks. Methylome data indicate that there were 3,307 and 3,776 differentially methylated regions (DMRs) with high-level methylation as well as 3,125 and 3,016 DMRs with low-level methylation in big papillae compared to small papillae. The identified DMRs were mainly distributed in introns, promotors, or exons. The whole-transcriptome and DNA methylome datasets generated from this study not only established a robust theoretical foundation (especially from the epigenetic aspect) for elucidating expression regulation network determining papilla size in sea cucumbers but also can be a valuable resource of biomarker mining for papilla appearance-based selective breeding in sea cucumbers.

DNA Methylation↗

The interobserver reliability and validity of volume calculation from three-dimensional ultrasound datasets in the in vitro setting.

OBJECTIVES: The primary aim of this validation study was to determine the interobserver reliability and validity of measurements of phantom objects of known volume using conventional and rotational techniques of volume calculation according to measurement technique. METHODS: Two observers each acquired a single three-dimensional ultrasound dataset of three water-filled objects of different size and shape. The same two observers measured all six datasets using both the conventional technique and the newer rotational technique (Virtual Organ Computer-aided AnaLysis, VOCAL( trade mark )) of volume calculation. Reliability was assessed by calculating intraclass correlation coefficients (ICC) and validity by examining the percentage difference from the 'true' volume, as determined by a water displacement technique, by the limits of agreement method. RESULTS: All of the techniques were highly reliable (ICC: 0.9962-0.9997) and valid to within 4% of the 'true' volumes. There were no significant differences in reliability according to measurement plane or between observers. Measurements made with the 6 degrees rotation step were significantly more reliable than those made by all other techniques with the exception of the 9 degrees rotation step (P < 0.05) and significantly more valid than those made with the 30 degrees rotation step or conventional technique (P < 0.05). CONCLUSIONS: Volume calculation in the in vitro setting is both reliable and valid but is dependent upon the technique applied, with rotational measurements of volume proving superior to conventional techniques.

Female↗

Phylogenetic analysis of a dataset of fungal 5.8S rDNA sequences shows that highly divergent copies of internal transcribed spacers reported from Scutellospora castanea are of ascomycete origin.

Using a dataset comprising 5.8S rDNA sequences from a wide range of fungi, we show that some sequences reported recently from the arbuscular mycorrhizal (AM) fungus Scutellospora castanea most likely originate from Ascomycetes. Other ITS and 5.8S sequences which were previously reported are confirmed as being clearly of mycorrhizal origin and are variable within one isolate of S. castanea. However, these results mean that previous conclusions which were drawn regarding the heterokaryotic status of AM fungal spores remain unproven. We provide an enlarged 5.8S rDNA dataset that can be used to check ITS sequences for conflicts with well-established phylogenies of the organisms that they were obtained from.

Amino Acid Sequence↗

A dataset of estimated heterozygous individual and carrier couple frequencies for pan-ancestry carrier screening.

The data described in this publication supported the development and evaluation of pan-ancestry reproductive carrier screening panels for autosomal recessive (AR) and X-linked (XL) conditions. Raw data included combined sets of DNA variants in 1,350 AR/XL genes obtained from the ClinVar and gnomAD databases. The dataset enabled calculations of positive yield for individuals and couples across both ancestry-specific and pan-ancestry, optimised "Goldilocks"-ranked gene panels, addressing population-specific variations in the frequencies of heterozygous individuals and carrier couples. The positive yield analysis offered a performance metric for carrier screening panels, facilitating the modeling of screening performance for panels of varying sizes and composition and providing resources for optimizing panel content to ensure equity across underrepresented genetic ancestries The dataset can support ongoing research into the equitable application of carrier screening and offers significant reuse potential for refining population genetic screening practices, validating computational models, and developing frameworks to update carrier screening panels in alignment with evolving genomic data, including in underrepresented and minority populations.

Carrier screening↗

Individual Differences in Cognitive Aging Rodent Datasets (ID-CARD): A collaborative platform for behavioral analysis across the lifespan.

Understanding cognitive aging requires approaches that capture individual variability while enabling integration across studies. In rodent models, behavioral data are central to this effort, yet cross-laboratory differences in experimental design limit comparability and constrain secondary analysis. To address this gap, we developed the Individual Differences in Cognitive Aging Rodent Datasets (ID-CARD), a first-of-its-kind collaborative repository aggregating trial-level Morris water maze data from multiple laboratories. ID-CARD is designed to support large-scale, integrative analyses and to facilitate secondary use of existing behavioral data in alignment with emerging data-sharing and transparency initiatives. Rather than imposing retrospective harmonization of experimental protocols, we implemented a normalization and modeling framework that enables comparison of learning trajectories while preserving meaningful variation across studies. Behavioral data from >&#x202f;5000 rats spanning common strains, both sexes, and multiple ages were normalized in training and performance domains and fit with a logarithmic function to derive an error accumulation rate coefficient (EARC) as a measure of spatial learning. Age was strongly associated with increased EARC, indicating attenuated learning, even after adjusting for non-spatial cue performance. Analyses of goodness of fit revealed systematic structure in learning dynamics, where age was associated with reduced learning-curve conformity after accounting for overall performance. Inter-individual variability in spatial learning also increased with age, with strain-specific interactions. These findings demonstrate that integrated analysis of heterogeneous behavioral datasets can yield robust, individual-level insights into cognitive aging. ID-CARD provides a scalable resource and analytic framework to advance discovery in behavioral neuroscience by enabling reuse, integration, and comparative analysis of existing data.

Cognitive aging↗

Three-dimensional computed tomography bronchoscopy using clinical datasets: a comparison with fibreoptic bronchoscopy.

OBJECTIVE: To assess three-dimensional computed tomography 'bronchoscopic' (3-DCTB) reconstruction of routine CT data as a non-invasive method of airway visualization, and compare it with fibreoptic bronchoscopy (FOB). METHODS: Fourteen datasets were acquired from 13 patients undergoing both FOB and CT examination of the chest. Standard continuous volume CT using 6 mm collimation and clinical FOB techniques were employed. Images were obtained from 3-DCTB reconstructions by segmentation and surface recognition algorithms generating surface rendered 'bronchoscopic views'. These were scored for technical quality and anatomical detail. The most distal bronchi seen in left upper and right lower lobes were recorded for FOB and 3-DCTB. RESULTS: On FOB, the subsegmental bronchi were seen in right lower and in left upper lobe in 10/14 cases and 4/14 cases, respectively. Visualization of the subsegmental airways was not achieved with 3-DCTB, as they could not be identified with confidence. 3-DCTB never achieved a more distal view than obtained by FOB. Using 3-DCT, the right, lower lobe segmental bronchi were seen in 10/14 cases, and lobar bronchus in 14/14 cases (two occluded). In the left upper lobe, 3-DCT showed segmental bronchi in 6/14 cases, lobar bronchus in 11/14 cases (one occluded) and the left main bronchus appeared occluded in 3/14 cases. Overall, technical quality and anatomical detail scores of the carina and proximal bronchi ranked significantly higher than views of segmental bronchi. CONCLUSIONS: 3-DCTB cannot routinely replace FOB for inspection of major and segmental bronchi. Subsegmental bronchi cannot be adequately demonstrated by 3-DCTB using 6 mm collimation datasets.

Adult↗

Time factors in larynx tumor radiotherapy: lag times and intertumor heterogeneity in clinical datasets from four centers.

PURPOSE: To use the time-dependent linear-quadratic model, both in the standard form and in a form modified to incorporate intertumor heterogeneity, in a reanalysis of 4 datasets for larynx tumor control, to provide more representative and direct estimates of the lag period, the time factor (lambda/alpha), and the clonogen population inactivation dose ([lnk]/alpha). METHODS AND MATERIALS: The data comprised 2,225 patients treated in Edinburgh (UK), Glasgow (UK), Manchester (UK), or Toronto (Canada), with tumor control assessed after at least 2 years. Heterogeneity in each series was taken into account using the coefficient of variation (CV) of the clonogen radiosensitivity (alpha). Maximum likelihood techniques were used to provide best estimates of the parameters, and also direct estimation of the more stable parameter ratios of interest. RESULTS: The use of different heterogeneity factors for the different series allowed common dose/time parameters to be fitted across all four series in a way not possible using the standard model, enabling the inherent effect of heterogeneity in flattening dose-response curves and in reducing time factors to be separated from the underlying more-representative values. Radiosensitivity CVs were calculated to be 30% (Edinburgh), 36% (Glasgow), 40% (Manchester), and 71% (Toronto). The lag phase was 32 days (95% CL 20-38 days) which was longer than the value of 23 days (11-36 days) deduced using the standard model without the heterogeneity parameter. The time factor was 1.2 (0.8-2.2) Gy/day, again greater than the value of 0.80 (0.54-1.41) Gy/day derived using the standard model. Similar larger time factors and longer lag periods could be reproduced using the standard model either by using a parameterization based on parameter ratios, or by omitting the discordant Toronto data and refitting just the data from the three UK centers. CONCLUSION: It was concluded that the heterogeneity model provides a better representation of the time factor for tumor control when data are analyzed comprising different stages of disease treated at different centers. The model allows different amounts of heterogeneity in different series, which tend to flatten dose-responses curves and reduce time factors, to be taken in to account. Also, direct maximum likelihood estimates can be made of the lag period, the time factor (lambda/alpha), and the fractionation sensitivity (beta/alpha), as well as the clonogen population inactivation dose (lnk)/alpha. Values of these parameter ratios are more robust and stable than the individual parameter values. The results of the present analysis using a total of 2,225 patients from four centers indicate that the average lag period may be somewhat longer and the average time factor somewhat greater (and the 95% confidence limits of the time factor exclude previous estimates), than the values deduced previously using simpler models and more diverse multi-center datasets.

Dose-Response Relationship, Radiation↗

Objective stenosis quantification from post-stenotic signal loss in phase-contrast magnetic resonance angiographic datasets of flow phantoms and renal arteries.

In this study a semi-automated and observer-independent algorithm for quantifying post-stenotic signal loss (PSL) in three-dimensional phase-contrast (PC) magnetic resonance angiography (MRA) of patients with renal artery stenosis is presented. This algorithm was developed on MRA datasets of stenotic phantoms, included in a flow circuit with stationary flows. The length and the severity of the PSL (incorporating both the length and the degree of PSL) in the MRA datasets were proposed for quantifying the stenoses. The algorithm was tested in renal arteries; ten patients with renal artery stenosis and seven healthy volunteers were investigated. Digital subtraction angiography was performed in the patients and served as the gold standard. Stenosis severity showed better correlation with the severity of the PSL than with the length, both for in vitro and in vivo measurements. Spearman correlation coefficients (rs) showed statistically significant correlations between the severity of the PSL and parameters determined by digital subtraction angiography, i.e., percent diameter stenosis (rs = 0.90). The length of the PSL showed no correlation with the diameter stenosis (rs = 0.37). In conclusion, this study presents a semi-automated and observer-independent way of quantifying signal loss, and the severity of the PSL is proposed for quantifying stenoses, rather than the length of PSL.

Aged↗

Virtual endoscopy: development and evaluation using the Visible Human datasets.

Virtual endoscopy (VE) is a new method of diagnosis using computer processing of 3D image datasets (such as CT or MRI scans) to provide simulated visualizations of patient specific organs similar or equivalent to those produced by standard endoscopic procedures. Conventional endoscopy is invasive and often uncomfortable for patients. It sometimes has serious side effects such as perforation, infection and hemorrhage. VE visualization avoids these risks and can minimize difficulties and decrease morbidity when used before actual endoscopic procedures. In addition, there are many body regions not compatible with real endoscopy that can be explored with VE. Eventually, VE may replace many forms of real endoscopy. There remains a critical need to refine and validate VE visualizations for routine clinical use. We have used the Visible Human Dataset from the National Library of Medicine to develop and test these procedures and to evaluate their use in a variety of clinical applications. We have developed specific clinical protocols to compare virtual endoscopy with real endoscopy. We have developed informative and dynamic on-screen navigation guides to help the surgeon or physician interactively determine body orientation and precise anatomical localization while performing the VE procedures. Additionally, the adjunctive value of full 3D imaging (e.g. looking "outside" of the normal field of view) during the VE exam is being evaluated. Quantitative analyses of local geometric and densitometric properties obtained from the virtual procedures ("virtual biopsy") are being developed and compared with other direct measures. Preliminary results suggest that these virtual procedures can provide accurate, reproducible and clinically useful visualizations and measurements. These studies will help drive improvements in and lend credibility to VE procedures and simulations as routine clinical tools. VE holds significant promise for optimizing endoscopic diagnostic procedures, minimizing patient risk and morbidity, and reducing health care costs.

Anatomy, Cross-Sectional↗

Using dependency/association rules to find indications for computed tomography in a head trauma dataset.

Analysis of a clinical head trauma dataset was aided by the use of a new, binary-based data mining technique, termed Boolean analyzer (BA), which finds dependency/association rules. With initial guidance from a domain user or domain expert, the BA algorithm is given one or more metrics to partition the entire dataset. The weighted rules are in the form of Boolean expressions. To augment the analysis of the rules produced, we applied a probabilistic interestingness measure (PIM) to order the generated rules based on event dependency, where events are combinations of primed and unprimed variables. Interpretation of the dependency rules generated on the clinical head trauma data resulted in a set of criteria that identified minor head trauma patients needing computed tomography (CT) scans. The BA criteria contained fewer variables than were found using recursive partitioning of Chi-square values (five variables versus seven variables, respectively). The BA five-variable criteria set was more sensitive but less specific than the seven-variable Chi-square criteria set. We believe that the BA method has broad applicability in the medical domain, and hope that this paper will stimulate other creative applications of the technique.

Algorithms↗