Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43Linked to original sources

A score test for determining sample size in matched case-control studies with categorical exposure.

The paper considers the problem of determining the number of matched sets in 1 : M matched case-control studies with a categorical exposure having k + 1 categories, k > or = 1. The basic interest lies in constructing a test statistic to test whether the exposure is associated with the disease. Estimates of the k odds ratios for 1 : M matched case-control studies with dichotomous exposure and for 1 : 1 matched case-control studies with exposure at several levels are presented in Breslow and Day (1980), but results holding in full generality were not available so far. We propose a score test for testing the hypothesis of no association between disease and the polychotomous exposure. We exploit the power function of this test statistic to calculate the required number of matched sets to detect specific departures from the null hypothesis of no association. We also consider the situation when there is a natural ordering among the levels of the exposure variable. For ordinal exposure variables, we propose a test for detecting trend in disease risk with increasing levels of the exposure variable. Our methods are illustrated with two datasets, one is a real dataset on colorectal cancer in rats and the other a simulated dataset for studying disease-gene association.

Algorithms↗

Contribution of morphometry in the differential diagnosis of fine-needle thyroid aspirates.

BACKGROUND: Cytologic discrimination of cellular nodules, follicular adenoma, and follicular carcinoma in the thyroid is problematic. Methods are needed to achieve a reliable diagnosis. Some sophisticated tools, such as microarrays, offer great potential but lack accompanying morphologic information. METHODS: One hundred twelve samples obtained from patients with lesions histopathologically diagnosed as nodular goiter, follicular adenoma, follicular carcinoma, and papillary carcinoma were used. Eight geometric features, such as nuclear area and circular form factor, were measured. The dataset was divided into six overlapping groups to represent the frequently encountered situations in routine practice. Multivariate analysis of variance, Tukey's honestly significant differences test, and discriminant analysis were performed. Statistical analysis was carried out with two conceptually different approaches. In the first, data from all measured nuclei were used. In the second, a subset of data representing the most extreme values of variables was extracted from the entire dataset to simulate the "selection procedure" performed during conventional morphologic examination. RESULTS: When the selected dataset instead of data from all measured nuclei was used, the correct classification rates in discriminant analysis improved considerably. CONCLUSIONS: Morphologic examination is based primarily on selection. Using data obtained from all of the cells in morphometry may cause a dilution effect in diagnostically important features. Morphometric studies may also be planned with a proper selection "bias." This may be particularly helpful when isolated abnormal cells carry most of the diagnostic information.

Adenocarcinoma, Follicular↗

An integrative approach to CTL epitope prediction: a combined algorithm integrating MHC class I binding, TAP transport efficiency, and proteasomal cleavage predictions.

Reverse immunogenetic approaches attempt to optimize the selection of candidate epitopes, and thus minimize the experimental effort needed to identify new epitopes. When predicting cytotoxic T cell epitopes, the main focus has been on the highly specific MHC class I binding event. Methods have also been developed for predicting the antigen-processing steps preceding MHC class I binding, including proteasomal cleavage and transporter associated with antigen processing (TAP) transport efficiency. Here, we use a dataset obtained from the SYFPEITHI database to show that a method integrating predictions of MHC class I binding affinity, TAP transport efficiency, and C-terminal proteasomal cleavage outperforms any of the individual methods. Using an independent evaluation dataset of HIV epitopes from the Los Alamos database, the validity of the integrated method is confirmed. The performance of the integrated method is found to be significantly higher than that of the two publicly available prediction methods BIMAS and SYFPEITHI. To identify 85% of the epitopes in the HIV dataset, 9% and 10% of all possible nonamers in the HIV proteins must be tested when using the BIMAS and SYFPEITHI methods, respectively, for the selection of candidate epitopes. This number is reduced to 7% when using the integrated method. In practical terms, this means that the experimental effort needed to identify an epitope in a hypothetical protein with 85% probability is reduced by 20-30% when using the integrated method. The method is available at http://www.cbs.dtu.dk/services/NetCTL. Supplementary material is available at http://www.cbs.dtu.dk/suppl/immunology/CTL.php.

ATP-Binding Cassette Transporters↗

Prediction of chromosomal aneuploidy from gene expression data.

Chromosomal aneuploidy is commonly observed in neoplastic diseases and is an important prognostic marker. Here we examine how gene expression profiles reflect aneuploidy and whether these profiles can be used to detect changes in chromosome copy number. We developed two methods for detecting such changes in the gene expression profile of a single sample. The first method, fold-change analysis, relies on the availability of gene expression data from a large cohort of patients with the same disease. The expression profile of the sample is compared with that of the dataset. The second method, chromosomal relative expression analysis, is more general and requires the expression data from the tested sample only. We found that the relative expression values are stable among different chromosomes and exhibit little variation between different normal tissues. We exploited this novel finding to establish the set of reference values needed to detect changes in the copy number of chromosomes in a single sample on the basis of gene expression levels. We measured the accuracy of the performance of each method by applying them to two independent leukemia datasets. The second method was also applied to two solid tumor datasets. We conclude that chromosomal aneuploidy can be detected and predicted by analysis of gene expression profiles. This article contains Supplementary Material available at http://www.interscience.wiley.com/jpages/1045-2257/suppmat.

Aneuploidy↗

A sparse marker extension tree algorithm for selecting the best set of haplotype tagging single nucleotide polymorphisms.

Single nucleotide polymorphisms (SNPs) play a central role in the identification of susceptibility genes for common diseases. Recent empirical studies on human genome have revealed block-like structures, and each block contains a set of haplotype tagging SNPs (htSNPs) that capture a large fraction of the haplotype diversity. Herein, we present an innovative sparse marker extension tree (SMET) algorithm to select optimal htSNP set(s). SMET reduces the search space considerably (compared to full enumeration strategy), and therefore improves computing efficiency. We tested this algorithm on several datasets at three different genomic scales: (1) gene-wide (NOS3, CRP, IL6 PPARA, and TNF), (2) region-wide (a Whitehead Institute inflammatory bowel disease dataset and a UK Graves' disease dataset), and (3) chromosome-wide (chromosome 22) levels. SMET offers geneticists with greater flexibilities in SNP tagging than lossless methods with adjustable haplotype diversity coverage (phi). In simulation studies, we found that (1) an initial sample size of 50 individuals (100 chromosomes) or more is needed for htSNP selection; (2) the SNP tagging strategy is considerably more efficient when the underlying block structure is taken into account; and (3) htSNP sets at 80-90% phi are more cost-effective than the lossless sets in term of relative power, relative risk ratio estimation, and genotyping efforts. Our study suggests that the novel SMET algorithm is a valuable tool for association tests.

Algorithms↗

Evaluation of PCA and ICA of simulated ERPs: Promax vs. Infomax rotations.

Independent components analysis (ICA) and principal components analysis (PCA) are methods used to analyze event-related potential (ERP) and functional imaging (fMRI) data. In the present study, ICA and PCA were directly compared by applying them to simulated ERP datasets. Specifically, PCA was used to generate a subspace of the dataset followed by the application of PCA Promax or ICA Infomax rotations. The simulated datasets were composed of real background EEG activity plus two ERP simulated components. The results suggest that Promax is most effective for temporal analysis, whereas Infomax is most effective for spatial analysis. Failed analyses were examined and used to devise potential diagnostic strategies for both rotations. Finally, the results also showed that decomposition of subject averages yield better results than of grand averages across subjects.

Algorithms↗

Pooled analysis of 3 European case-control studies of ovarian cancer: II. Age at menarche and at menopause.

The role of age at menarche and at menopause on epithelial ovarian cancer risk was re-assessed in a combined analysis of 3 hospital-based case-control studies conducted in Italy, the United Kingdom and Greece, which produced a total of 1,140 cases and 2,724 controls. In the overall dataset, there was no evidence of an association with age at menarche: compared with women whose menarche occurred at age 15 or over, the relative risk (RR) estimates were 1.0 [95% confidence interval (CI) 0.8 to 1.2] for those with menarche at ages 12 to 14, and 1.0 (95% CI 0.8-1.2) for those with menarche below age 12. There was no consistent interaction between age at menarche and study centre or age at diagnosis. In relation to age at menopause, compared with women whose menopause occurred at age 44 or earlier, the RR was 1.4 between 45 and 49, 1.6 between 50 and 52 and 1.9 above 52. The strength of the association was apparently (but not significantly) greater in the British than in the Greek or Italian dataset. The effect of age at menopause tended to be long-lasting and, if anything, to increase across subsequent age-groups. The large dataset, and the replication of results in different studies, provide more definite and precise information than previously available on the absence of appreciable effect of age at menarche on subsequent ovarian cancer risk in developed countries. For age at menopause, there was a direct and consistent trend in risk, but the association was relatively weak, with RRs below 2 even between extreme categories.

Adult↗

Prediction and classification of protein subcellular location-sequence-order effect and pseudo amino acid composition.

Given a protein sequence, how to identify its subcellular location? With the rapid increase in newly found protein sequences entering into databanks, the problem has become more and more important because the function of a protein is closely correlated with its localization. To practically deal with the challenge, a dataset has been established that allows the identification performed among the following 14 subcellular locations: (1) cell wall, (2) centriole, (3) chloroplast, (4) cytoplasm, (5) cytoskeleton, (6) endoplasmic reticulum, (7) extracellular, (8) Golgi apparatus, (9) lysosome, (10) mitochondria, (11) nucleus, (12) peroxisome, (13) plasma membrane, and (14) vacuole. Compared with the datasets constructed by the previous investigators, the current one represents the largest in the scope of localizations covered, and hence many proteins which were totally out of picture in the previous treatments, can now be investigated. Meanwhile, to enhance the potential and flexibility in taking into account the sequence-order effect, the series-mode pseudo-amino-acid-composition has been introduced as a representation for a protein. High success rates are obtained by the re-substitution test, jackknife test, and independent dataset test, respectively. It is anticipated that the current automated method can be developed to a high throughput tool for practical usage in both basic research and pharmaceutical industry.

Algorithms↗

A vision-based, 3D reconstruction technique for scanning electron microscopy: direct comparison with atomic force microscopy.

High-resolution, detailed 3D reconstructions of biological specimens obtained from scanning electron microscopy stereo-micrographs and proprietary software were compared with Tapping-Mode AFM datasets of the same fields. The reconstruction software implements several original solutions including a neural adaptive point-matching technique, the ability to build an irregular triangulated mesh rather than a regular orthogonal grid, and the ability to re-map one of the original images exactly onto the reconstructed surface. The technique was applied to human nerve tissue to obtain 1,424 x 968-pixel, texture-mapped datasets, which were subsequently compared against 512 x 512-pixel AFM datasets from the same viewfields. Accounting for the inherent differences of the two techniques, direct comparison revealed an excellent visual match. The correspondence was also quantified by calculating the cross-correlation coefficient between corresponding altimetric profiles in SEM and AFM data, which consistently exceeded a figure of 0.9, with a rate of point mismatch in the order of 0.01%. Research is still underway to improve the robustness of the technique when applied to arbitrary images

Humans↗

Coil-by-coil image reconstruction with SMASH.

The SiMultaneous Acquisition of Spatial Harmonics (SMASH) technique uses linear combinations of undersampled datasets from the component coils of an RF coil array to reconstruct fully sampled composite datasets in reduced imaging times. In previously reported implementations, SMASH reconstructions were designed to reproduce the images that would otherwise be obtained by simple sums of fully gradient encoded component coil images. This strategy has left SMASH images vulnerable to phase cancellation artifacts when the sensitivities of RF coil array elements are not suitably phase-aligned. In fully gradient encoded imaging schemes these artifacts can be eliminated using a variety of methods for combining the individual coil images, including matched filter combinations as well as sum of squares combinations. Until now, these reconstruction schemes have been unavailable to SMASH reconstructions as SMASH produced a final composite image directly from the raw component coil k-space datasets. This article demonstrates a modification to SMASH that allows reconstruction of a full set of accelerated individual component coil images by fitting component coil sensitivity functions to a complete set of spatial harmonics tailored for each coil in the array. Standard component coil combinations applied to the individual reconstructed images produce final composite images free of phase cancellation artifacts.

Abdomen↗

Ovarian cancer detection by logical analysis of proteomic data.

A new type of efficient and accurate proteomic ovarian cancer diagnosis systems is proposed. The system is developed using the combinatorics and optimization-based methodology of logical analysis of data (LAD) to the Ovarian Dataset 8-7-02 (http://clinicalproteomics.steem.com), which updates the one used by Petricoin et al. in The Lancet 2002, 359, 572-577. This mass spectroscopy-generated dataset contains expression profiles of 15 154 peptides defined by their mass/charge ratios (m/z) in serum of 162 ovarian cancer and 91 control cases. Several fully reproducible models using only 7-9 of the 15 154 peptides were constructed, and shown in multiple cross-validation tests (k-folding and leave-one-out) to provide sensitivities and specificities of up to 100%. A special diagnostic system for stage I ovarian cancer patients is shown to have similarly high accuracy. Other results: (i) expressions of peptides with relatively low m/z values in the dataset are shown to be better at distinguishing ovarian cancer cases from controls than those with higher m/z values; (ii) two large groups of patients with a high degree of similarities among their formal (mathematical) profiles are detected; (iii) several peptides with a blocking or promoting effect on ovarian cancer are identified.

Algorithms↗

Predicting protein secondary structure with probabilistic schemata of evolutionarily derived information.

We demonstrate the applicability of our previously developed Bayesian probabilistic approach for predicting residue solvent accessibility to the problem of predicting secondary structure. Using only single-sequence data, this method achieves a three-state accuracy of 67% over a database of 473 non-homologous proteins. This approach is more amenable to inspection and less likely to overlearn specifics of a dataset than "black box" methods such as neural networks. It is also conceptually simpler and less computationally costly. We also introduce a novel method for representing and incorporating multiple-sequence alignment information within the prediction algorithm, achieving 72% accuracy over a dataset of 304 non-homologous proteins. This is accomplished by creating a statistical model of the evolutionarily derived correlations between patterns of amino acid substitution and local protein structure. This model consists of parameter vectors, termed "substitution schemata," which probabilistically encode the structure-based heterogeneity in the distributions of amino acid substitutions found in alignments of homologous proteins. The model is optimized for structure prediction by maximizing the mutual information between the set of schemata and the database of secondary structures. Unlike "expert heuristic" methods, this approach has been demonstrated to work well over large datasets. Unlike the opaque neural network algorithms, this approach is physicochemically intelligible. Moreover, the model optimization procedure, the formalism for predicting one-dimensional structural features and our previously developed method for tertiary structure recognition all share a common Bayesian probabilistic basis. This consistency starkly contrasts with the hybrid and ad hoc nature of methods that have dominated this field in recent years.

Algorithms↗

Variability of pesticide residues in crop units.

The results of 89 new field trials and 11 supervised trials were considered, together with 91 sets of residue data evaluated earlier. The datasets consisted of 22,643 valid residue data. As all variability factors calculated from individual sample sets are affected by the uncertainties of sampling and analysis, the average of the P(0.975)/R(ave) (97.5th percentile of the residue population divided by the average residues in the lot) values gives the best estimate for the variability factor. The Harrell-Davis (H-D) method gave an average value of 2.89 for the variability factor for all samples, while the average variability factors obtained from samples derived from the new field and supervised trials were 2.8 and 2.7 with the IUPAC and H-D methods respectively. The number of residue values below the LOQ in a sample set significantly affects the observed variability factors. It was found that datasets containing over 20% non-quantifiable residues might not reflect the true variability of the residues. Mixing of treated and non-treated commodities may significantly increase the apparent variability. Consequently, only datasets of known origin and consisting of well-quantifiable residues should be used for estimation of the variability factor. Samples taken from marketed lots may not represent a single lot, and thus they have limited value in estimating the variability factor. The large number of residue data confirms the applicability of the default variability factor of 3 adopted by the FAO/WHO for deterministic estimation of the acute intake of pesticide residues.

Fruit↗

Analysis of repeated categorical data using generalized estimating equations.

Moment methods for analysing repeated binary responses have been proposed by Liang and Zeger, and extended by Prentice and Zhao and Prentice. In these estimating equations, models are proposed for the correlation between the repeated binary responses. We extend Liang and Zeger's method to models for the correlation between repeated nominal or ordinal categorical responses; in particular, when the repeated responses are binary, our methods reduce to Liang and Zeger's method. Our method is illustrated with two datasets. One dataset contains repeated observations of self-assessment of arthritis, an ordered variable with three categories, collected during a randomized comparative study of alternative treatments of patients with rheumatoid arthritis. The second dataset is a longitudinal study of the health effects of air pollution, in which the repeated ordered multinomial response is the wheezing status (no wheeze, wheeze with cold, wheeze apart from cold) of a child at ages 9, 10, 11 and 12 years.

Adult↗

The 'spin' technique: a new method for examination of the fetal outflow tracts using three-dimensional ultrasound.

BACKGROUND AND OBJECTIVE: The prenatal detection of congenital heart defects remains one of the most difficult challenges for the sonologist/sonographer when performing the second- or third-trimester screening examination. The four-chamber view has been used for a number of years as the primary screening image for detection of heart defects, but the inclusion of the right and left outflow tracts increases the detection of cardiac malformations. One of the difficulties, however, is obtaining and interpreting two-dimensional images of the outflow tracts. This paper reviews a new technique using three-dimensional (3D) multiplanar imaging that allows the examiner to identify the outflow tracts within a few minutes of acquiring the 3D volume dataset by rotating the volume dataset around the x- and y-axes. METHODS: 3D multiplanar imaging of the fetal heart using static 3D or spatio-temporal image correlation (STIC) imaging allows the examiner to obtain a volume of data that can be manipulated along the x- and y-axes using reference points from the four-chamber view, five-chamber view, three-vessel view at the level of the bifurcation of the pulmonary arteries, and three-vessel view at the level of the transverse aortic arch and trachea. RESULTS: The full length of the main pulmonary artery, ductus arteriosus, aortic arch and superior vena cava could be identified easily in the normal fetus by rotating the volume dataset along the x- and y-axes. The vessels were identified using the four-chamber view, the five-chamber view, and the two three-vessel views. The technique was useful in identification of d-transposition of the great vessels and evaluation of the outflow tracts in hypoplastic left heart syndrome. CONCLUSION: 3D multiplanar evaluation of the fetal heart allows the examiner to identify the outflow tracts using a simple technique that requires only rotation around x- and y-axes from reference images obtained in a transverse sweep through the fetal chest.

Aorta, Thoracic↗

Reliability and validity of tissue volume measurement by three-dimensional ultrasound: an experimental model.

OBJECTIVE: To determine the validity and the intra- and interobserver reliability of volume measurements of an endometrium-like model using a three-dimensional (3D) ultrasound rotational technique. METHODS: A 3D ultrasound dataset was obtained from a sample of bovine liver containing a portion of chicken chest muscle (CCM). The process was repeated seven times using pieces of CCM of different sizes, resulting in seven datasets. Each portion of CCM was then placed in a water-filled volume-scaled tube and the 'actual' volumes were calculated by water displacement. For each dataset, ten volumes were calculated by each of two observers using a (VOCAL) with a 15 degrees rotational step. Reliability was assessed by calculating intraclass correlation coefficients (ICC) and validity by examining the percentage difference from the actual volume using limits of agreement. RESULTS: The volume measurement of organic tissues using the 3D ultrasound rotational method was highly reliable (intraobserver ICC, 0.998 for Observer 1 and 0.997 for Observer 2; interobserver ICC, 0.997) and valid (the bias and 95% limits of agreement of the percentage difference from the actual volume was only 0.57 (-3.07 to 4.21) % for Observer 1 and - 0.17 (-4.34 to 4.0) % for Observer 2). CONCLUSIONS: The 3D sonographic measurement, using VOCAL with a 15 degrees rotational step, of small and irregular tissues is reliable and valid, suggesting that it is a useful technique for measurement of the endometrial volume and other volumes of similar size.

Animals↗

Exploratory visualization software for reporting environmental survey results.

Environmental surveys yield three principal products: maps, a set of data tables, and a textual report. The relationships between these three elements, however, are often cumbersome to present, making full use of all the information in an integrated and systematic sense difficult. The published paper report is only a partial solution. Modern developments in computing, particularly in cartography, GIS, and hypertext, mean that it is increasingly possible to conceive of an easier and more interactive approach to the presentation of such survey results. Here, we present such an approach which links map and tabular datasets arising from a vegetation survey, allowing users ready access to a complex dataset using dynamic mapping techniques. Multimedia datasets equipped with software like this provide an exciting means of quick and easy visual data exploration and comparison. These techniques are gaining popularity across the sciences as scientists and decision-makers are presented with increasing amounts of diverse digital data. We believe that the software environment actively encourages users to make complex interrogations of the survey information, providing a new vehicle for the reader of an environmental survey report.

Conservation of Natural Resources↗

Hamiltonians for protein tertiary structure prediction based on three-dimensional environment principles.

We describe a computational approach to protein tertiary structure prediction that combines ideas from the three-dimensional (3D) profile method of Bowie, Lüthy and Eisenberg and the associative memory Hamiltonians of Friedrichs and Wolynes. The ultimate goal of our work is to extend and generalize the capabilities of these heuristics so as to be able to predict novel structures that might be found in nature or designed proteins. In our approach we approximate the interactions between residues through a pseudo-potential function similar to an associative memory Hamiltonian. This function is constructed based on 3D environment principles. Favorable inter-residue contacts for each residue in a target protein are inferred by using 3D environment propensities of the residues and a collection of 3D environment templates derived from a dataset of protein crystal structures. A Hamiltonian encoding this information is used to guide an optimization phase via molecular dynamics with annealing, which then leads to the folded structure. With our algorithm we can recover the structure of dataset proteins and have also succeeded in constructing the fold for a protein with little sequence similarity to any dataset protein.

Aprotinin↗