Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

[Estimation of disease-specific costs in a dataset of health insurance claims and its validation using simulation data].

PURPOSES: To estimate disease-specific costs in a dataset of health insurance claims with multiple diagnoses with known aggregate cost per claim and unknown disease-specific cost of each diagnosis using PDM (Proportional Disease Magnitude) method, validate its accuracy using simulation data with Monte Carlo method and improve its accuracy by developing an adjustment formula. METHODS: Developed simulation data with pre-assigned disease-specific costs, applied PDM method using arithmetic means of per-diem-per-disease cost as magnitude, validated its accuracy by observing the correlation between estimates by PDM method and known disease-specific costs and formulated an adjustment formula to improve accuracy. The reproducibility of the findings was assessed using Monte Carlo method by repeating the same procedures. RESULTS: The observed arithmetic means of per-diem-per-disease cost did not match well with actual values resulting in unsatisfactory accuracy. However, when the observed means were adjusted with a formula in which the observed mean is multiplied by (observed mean/overall mean) in the power of 2, PDM method yielded an accurate estimate of disease-specific cost. The accuracy was reproduced by Monte Carlo method with 0.9 or above R square value and slope of regression line in 76, 56 out of 100 iterations respectively. CONCLUSIONS: PDM method proved to be an objective, reproducible and accurate method for estimation of disease-specific costs of health insurance claims.

Insurance Claim Reporting↗

Using symbolic knowledge in the UMLS to disambiguate words in small datasets with a naïve Bayes classifier.

Current approaches to word sense disambiguation use and combine various machine-learning techniques. Most refer to characteristics of the ambiguous word and surrounding words and are based on hundreds of examples. Unfortunately, developing large training sets is time-consuming. We investigate the use of symbolic knowledge to augment machine-learning techniques for small datasets. UMLS semantic types assigned to concepts found in the sentence and relationships between these semantic types form the knowledge base. A naïve Bayes classifier was trained for 15 words with 100 examples for each. The most frequent sense of a word served as the baseline. The effect of increasingly accurate symbolic knowledge was evaluated in eight experimental conditions. Performance was measured by accuracy based on 10-fold cross-validation. The best condition used only the semantic types of the words in the sentence. Accuracy was then on average 10% higher than the baseline; however, it varied from 8% deterioration to 29% improvement. In a follow-up evaluation, we noted a trend that the best disambiguation was found for words that were the least troublesome to the human evaluators.

Abstracting and Indexing↗

A 3D interactive multimodal viewer as data mining tool for the Visible Human Dataset color image histograms.

An on-line virtual three-dimensional immersive environment to navigate through colorimetric characterization of the Visible Human Dataset (VHD) cryosectional cross-section color images is introduced. Real-time analysis of color component characteristics of a user defined set of VHD images is now possible. This is a potentially useful resource to many developers working on the VHD raw data, however it could be used in medical education.

Female↗

Large medical datasets on the Grid.

OBJECTIVE: This paper shows the use of the emerging Grid technology for gathering underused resources that are distributed among a corporate network. The work of these resources is coordinated for facing tasks which are not affordable by the individual usage of each of them. METHODS: This paper shows an application for the projection, using Volume Rendering techniques, of huge medical volumes obtained from CTs and RMIs, adapted to Grid computing. RESULTS: As a result the article shows the feasibility of the creation of an application based up on Grid technology, which solves problems that cannot be addressed by using common techniques. As an example, the article describes the projection of a huge medical dataset, which exceeds the resources of most common PCs, carried out by taking profit of idle CPU cycles from the computers of an organization. CONCLUSIONS: Grid technology is emerging as a new framework which allows gathering and coordinating resources distributed among a network (LAN or WAN), for addressing problems which cannot be solved through the single use of any of these resources. Medical Imaging is a clear application area for this technology.

Algorithms↗

Calibration and validation of multiple regression models for stormwater quality prediction: data partitioning, effect of dataset size and characteristics.

Two main issues regarding stormwater quality models have been investigated: i) the effect of calibration dataset size and characteristics on calibration and validation results; ii) the optimal split of available data into calibration and validation subsets. Data from 13 catchments have been used for three pollutants: BOD, COD and SS. Three multiple regression models were calibrated and validated. The use of different data sets and different models allows viewing general trends. It was found mainly that multiple regression models are case sensitive to calibration data. Few data used for calibration infers bad predictions despite good calibration results. It was also found that the random split of available data into halves for calibration and validation is not optimal. More data should be allocated to calibration. The proportion of data to be used for validation increases with the number of available data (N) and reaches about 35% for N around 55 measured events.

Calibration↗

Dilation based modeling of perfusion datasets.

A new approach to the modeling of the marker in Perfusion CT and Perfusion MR datasets is outlined and initial results given. The technique is based on estimation of the dilation and delay of an estimated bolus shape and a template fit to an new solution of the heat equation. Initial results are provided.

Brain↗

Improving the comparability of cancer registry treatment data and proposals for a new national minimum dataset.

BACKGROUND: There is no consistent or standardized practice for the collection of treatment data in UK cancer registries. This limits the usefulness and effectiveness of undertaking multiregional or national studies of treatment outcomes and survival. METHODS: A working group was established to examine the practices for recording the type and the amount of treatment data held in the cancer records at different registries. A common set of anonymized case notes for breast and colorectal cancer patients, drawn from each registry, was employed to eliminate any selection bias. Each registry coded these case notes according to their own criteria, and the comparability of such data between registries was determined from their returns. RESULTS: Of the 11 registries in England, seven participated in the full study, with a total of 84 records being submitted by five registries. A flow diagram was constructed to show how specific data items in the cancer record structure could be linked between registries. Errors or inconsistencies in recording treatment details were identified, and the constraints in data comparability were defined from the case note returns. CONCLUSION: Variations in coding practice between registries were such as to vitiate interregional or national comparisons of current data. The working group recommended an extended minimum dataset, which included a date for the start of each treatment modality, that most registries should be able to implement with some system changes.

Data Collection↗

PotatoRTD and TomatoRTD: Comprehensive Reference Transcript Datasets for Accurate Transcriptome Analysis and Isoform Discovery.

Transcriptome annotations provide essential information on transcript locations, sequences and structures, including transcription start, end sites and splice junctions. They underpin key biological analyses such as gene and transcript quantification, and the study of transcriptional and post-transcriptional regulation, including alternative transcription initiation, polyadenylation and splicing. Accurate characterisation of transcript isoforms is critical for understanding how gene expression relates to functional protein products. However, for many species-including Solanaceae crops such as potato and tomato-current annotations suffer from limited isoform coverage, with tens or hundreds of thousands of splice junctions and transcript isoforms missing. This undermines the completeness and accuracy of transcript-level analyses. Here, by generating Iso-seq and RNA-seq on a range of tissues and samples, we have produced transcriptome annotations for both potato and tomato with improved coverage, diversity, accurate splice junctions, and transcript start and end sites. We have also made these high-quality resources accessible through genome browsers. These enhanced annotations will enable more accurate transcriptome analyses, supporting higher-resolution and novel biological discoveries.

Solanum tuberosum↗

Genetic Ancestry and Colorectal Cancer in the All of Us Dataset.

IMPORTANCE: Genetic ancestry may complement biological, behavioral, and clinical factors in understanding colorectal cancer (CRC) disparities; yet, ancestry-informed analyses in CRC remain limited. OBJECTIVE: To characterize associations of genetic ancestry with CRC burden, age at diagnosis, and age-specific risk, and to develop a multiethnic CRC risk-prediction model. DESIGN, SETTING, AND PARTICIPANTS: This retrospective cohort study used All of Us data from July 1986 to October 2023, with follow-up through last visit or death (median [IQR], 133.1 [57.1-186.5] months); analyses were conducted from February to June 2026. All of Us is a US research cohort with linked electronic health record (EHR) and short-read whole-genome sequencing (srWGS) data. All of Us Research Program participants with srWGS and linked EHR data were included, except those with hereditary polyposis or Lynch syndrome. EXPOSURES: Genetically inferred ancestry categories and principal components. MAIN OUTCOMES AND MEASURES: Any CRC was the primary outcome. Associations were evaluated using Fisher exact tests, cumulative incidence functions with Gray tests, cause-specific and Fine-Gray subdistribution hazard models, and pooled multivariable logistic regression. Prediction models used penalized least absolute shrinkage and selection operator and extreme gradient boosting (XGBoost). RESULTS: Among 316 624 participants (median [IQR] age, 56.3 [40.2-68.2] years; 172 327 [54.4%] of European ancestry; 191 705 female [61.2%]; 121 585 male [38.8%]), 2914 (0.9%) developed CRC. European ancestry was associated with higher odds of CRC vs all other ancestries combined (odds ratio, 1.50; 95% CI, 1.39-1.62). The median age at CRC diagnosis was older in European (63.4 [53.9-71.2] years) than in American admixed-Latino, African, East Asian, and Other ancestry groups. In cause-specific hazard models on the attained-age scale, American admixed-Latino (hazard ratio, 1.30; 95% CI, 1.14-1.47) and East Asian (hazard ratio, 1.43; 95% CI, 1.06-1.94) ancestry had higher age-specific CRC hazard than European ancestry, with consistent findings on the subdistribution scale accounting for competing death. The multiethnic XGBoost model performed best (receiver operating characteristic area under the curve, 0.898; 95% CI, 0.882-0.912; precision-recall area under the curve, 0.338; 95% CI, 0.296-0.379) and was well calibrated. CONCLUSIONS AND RELEVANCE: In this cohort study, genetic ancestry was associated with meaningful differences in CRC burden and age-specific risk. These findings suggest that a multiethnic XGBoost model may complement CRC screening as a risk-enrichment tool.

Aged↗

The Visible Human Dataset: the anatomical platform for human simulation.

One goal of a medical school education is to teach the anatomy of the living human. With the exception of some surface anatomy, the morphology education that goes on during a surgical procedure, and patient observation, live human anatomy is most often taught by simulation. Medical anatomy courses utilize cadavers to approximate the live human. Case-based curricula simulate a patient and present symptoms, signs, and history to mimic reality for the future practitioner. Radiology has provided images of the morphology, function, and metabolism of living humans but with images foreign to most novice observers. With the Visible Human database, computer simulation of the live human body will provide revolutionary transformations in anatomical education.

Anatomy, Cross-Sectional↗

Localization of autosomal recessive early-onset parkinsonism to chromosome 1p36 (PARK7) in an independent dataset.

Two new loci, PARK6 and PARK7, for autosomal recessive early-onset parkinsonism have recently been identified on chromosome 1p, in single large pedigrees. Among 4 autosomal recessive early-onset families analyzed here, 2 supported linkage to PARK7, 1 with conclusive evidence. These data confirm localization of autosomal recessive early-onset parkinsonism to PARK7, suggesting it to be a frequent locus. Assignment of families to either PARK6 or PARK7 might be difficult because of the proximity of the two loci on chromosome 1p.

Adult↗

Prediction of aqueous solubility based on large datasets using several QSPR models utilizing topological structure representation.

Several QSPR models were developed for predicting intrinsic aqueous solubility, S(o). A data set of 5,964 neutral compounds was sub-divided into two classes, aromatic and non-aromatic compounds. Three models were created with different methods on both data sets: two regression models (multiple linear regression and partial least squares) and an artificial neural network model. These models were based on 3343 aromatic and 1674 non-aromatic compounds for training sets; 938 compounds were used in external validation testing. The range in -log S(o) is -1.6 to 10. Topological structure descriptors were used with all models. A genetic algorithm was used for descriptor selection for regression models. For the artificial neural network (ANN) model, descriptor selection was done with a backward elimination process. All models performed well with r2 values ranging 0.72 to 0.84 in external validation testing. The mean absolute errors in validation ranged from 0.44 to 0.80 for the classes of compounds for all the models. These statistical results indicate a sound ANN model. Furthermore, in a comparison with eight other available models, based on predictions using a validation test set (442 compounds), the artificial neural network model presented in this work (CSLogWS) was clearly superior based on both the mean absolute error and the percentage of residuals less than one log unit. In the ANN model both E-State and hydrogen E-State descriptors were found to be important.

Databases, Factual↗

HLA-DR effects in a large German IDDM dataset.

DR4 and DR3 are in strongest linkage disequilibrium with IDDM susceptibility genes, and DR1 demonstrates a lesser degree of positive disequilibrium. DR3/DR4 heterozygotes have the highest risk. The DR1 increase occurs almost exclusively in DR4/DR1 heterozygotes, suggesting that DR1 may be in disequilibrium with the same susceptibility gene as DR3. Homozygotes for DR4 and especially DR3 have a higher risk than heterozygotes of either with DRX. GLO-2 is increased in diabetic haplotypes carrying DR4, DR3 or DR1.

Alleles↗

A two-step procedure for constructing confidence intervals of trait loci with application to a rheumatoid arthritis dataset.

Preliminary genome screens are usually succeeded by fine mapping analyses focusing on the regions that signal linkage. It is advantageous to reduce the size of the regions where follow-up studies are performed, since this will help better tackle, among other things, the multiplicity adjustment issue associated with them. We describe a two-step approach that uses a confidence set inference procedure as a tool for intermediate mapping (between preliminary genome screening and fine mapping) to further localize disease loci. Apart from the usual Hardy-Weiberg and linkage equilibrium assumptions, the only other assumption of the proposed approach is that each region of interest houses at most one of the disease-contributing loci. Through a simulation study with several two-locus disease models, we demonstrate that our method can isolate the position of trait loci with high accuracy. Application of this two-step procedure to the data from the Arthritis Research Campaign National Repository also led to highly encouraging results. The method not only successfully localized a well-characterized trait contributing locus on chromosome 6, but also placed its position to narrower regions when compared to their LOD support interval counterparts based on the same data.

Arthritis, Rheumatoid↗

Accumulating quantitative trait linkage evidence across multiple datasets using the posterior probability of linkage.

Genome scans for complex disorders are frequently inconclusive, prompting researchers to increase sample size in an effort to obtain stronger evidence. However, increasing sample size in the presence of locus heterogeneity may actually, on average, decrease the linkage signal at a true susceptibility gene. The posterior probability of linkage, or PPL, was specifically designed to address this issue in the context of categorical trait analysis, by appropriately accumulating evidence either for or against linkage as new data are added. We now formulate a quantitative trait (QT) analog, the QT-PPL, which directly measures the evidence that a QT is linked to a genetic marker or location. The new QT-PPL is based on a classical single-locus QT likelihood with the trait parameters (allele frequency, genotypic means and variances) integrated out. We show using simulations that the QT-PPL is robust to two key modeling violations (multiple trait loci and non-normality in the form of excess kurtosis), as well as being inherently ascertainment corrected, and illustrate the advantages of the QT-PPL for accumulating linkage evidence across multiple sets of data compared to other QT linkage methods.

Computer Simulation↗

Dealing with the shortcomings of spatial normalization: multi-subject parcellation of fMRI datasets.

The analysis of functional magnetic resonance imaging (fMRI) data recorded on several subjects resorts to the so-called spatial normalization in a common reference space. This normalization is usually carried out on a voxel-by-voxel basis, assuming that after coregistration of the functional images with an anatomical template image in the Talairach reference system, a correct voxel-based inference can be carried out across subjects. Shortcomings of such approaches are often dealt with by spatially smoothing the data to increase the overlap between subject-specific activated regions. This procedure, however, cannot adapt to each anatomo-functional subject configuration. We introduce a novel technique for intra-subject parcellation based on spectral clustering that delineates homogeneous and connected regions. We also propose a hierarchical method to derive group parcels that are spatially coherent across subjects and functionally homogeneous. We show that we can obtain groups (or cliques) of parcels that well summarize inter-subject activations. We also show that the spatial relaxation embedded in our procedure improves the sensitivity of random-effect analysis.

Algorithms↗

Identifying residues in natural organic matter through spectral prediction and pattern matching of 2D NMR datasets.

This paper describes procedures for the generation of 2D NMR databases containing spectra predicted from chemical structures. These databases allow flexible searching via chemical structure, substructure or similarity of structure as well as spectral features. In this paper we use the biopolymer lignin as an example. Lignin is an important and relatively recalcitrant structural biopolymer present in the majority of plant biomass. We demonstrate how an accurate 2D NMR database of approximately 600 2D spectra of lignin fragments can be easily constructed, in approximately 2 days, and then subsequently show how some of these fragments can be identified in soil extracts through the use of various search tools and pattern recognition techniques. We demonstrate that once identified in one sample, similar residues are easily determined in other soil extracts. In theory, such an approach can be used for the analysis of any organic mixtures.

Benzopyrans↗