Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sampling Errors”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,135 records · Page 63Linked to original sources

Superior feature-set ranking for small samples using bolstered error estimation.

MOTIVATION: Ranking feature sets is a key issue for classification, for instance, phenotype classification based on gene expression. Since ranking is often based on error estimation, and error estimators suffer to differing degrees of imprecision in small-sample settings, it is important to choose a computationally feasible error estimator that yields good feature-set ranking. RESULTS: This paper examines the feature-ranking performance of several kinds of error estimators: resubstitution, cross-validation, bootstrap and bolstered error estimation. It does so for three classification rules: linear discriminant analysis, three-nearest-neighbor classification and classification trees. Two measures of performance are considered. One counts the number of the truly best feature sets appearing among the best feature sets discovered by the error estimator and the other computes the mean absolute error between the top ranks of the truly best feature sets and their ranks as given by the error estimator. Our results indicate that bolstering is superior to bootstrap, and bootstrap is better than cross-validation, for discovering top-performing feature sets for classification when using small samples. A key issue is that bolstered error estimation is tens of times faster than bootstrap, and faster than cross-validation, and is therefore feasible for feature-set ranking when the number of feature sets is extremely large.

Algorithms↗

"Staircase" saccadic intrusions plus transient yoking and neural integrator failure associated with cerebellar hypoplasia: a model simulation.

We present hypothesized ocular motor mechanisms of unique "staircase-like" sequences of saccadic intrusions in one direction that we have named, "staircase saccadic intrusions (SSI)," square-wave jerks/oscillations (SWJ/SWO), and transient failures of yoking and neural integrators in a patient with severe hypotonia, ataxic speech, motor and language developmental delays, and torticollis (Joubert syndrome). Brain magnetic resonance imaging showed hypoplasia of the cerebellar vermis and inferior cerebellar peduncles, abnormal superior cerebellar peduncles with deepening of the interpeduncular fossa, and enlargement of the fourth ventricle. During far and near fixation and smooth pursuit (rightward markedly better than leftward), the subject exhibited conjugate SSI (rightward more than leftward, with intersaccadic intervals equivalent to the normal 250 msec visual latency), SWJ, SWO, and uniocular, convergent and divergent saccades (including double saccades). Simulations using a behavioral ocular motor system model identified hypothetical mechanisms for SWJ, SWO, and SSI and ruled out the loss of efference copy as the cause. SSI may result from simultaneous dysfunctions: 1) a transient loss of accurate retinal-error information and/or sampled, reconstructed error; plus 2) a constant sampled, reconstructed retinal error that drives saccades.

Cerebellar Ataxia↗

Hydration (water binding) of the mammalian corneal stroma ex vivo and vitro: sample mass and error considerations.

The corneal stroma is well-known for its capacity to absorb water from experimental solutions but there are considerable differences in how this is actually assessed and described. Measurements on rabbit, sheep, and cow samples ex vivo are presented along with examples of the progressive changes in the stroma in vitro in saline solutions. These are used as the basis for theoretically assessing the impact of systematic and unintentional errors (related to excesses of water or water vapor on the samples) on the estimates of water content, as are the effects of unintentional loss of water from wet samples. Substantial differences in data can easily arise from such errors and/or differences in sample mass. Standardization of the methods for stromal hydration and swelling is required. As a minimum, sample origin (including species), sample dimensions and mass, and the resolution of the weighing protocols are needed, along with details of dry and wet sample handling. With the likelihood of modest errors for the commonly used samples and procedures, reporting hydration data to more than one decimal place is not justified.

Animals↗

Comparison of quantitative diagnostic tests: type I error, power, and sample size.

For a quantitative laboratory test the 0.975 fractile of the distribution of reference values is commonly used as a discrimination limit, and the sensitivity of the test is the proportion of diseased subjects with values exceeding this limit. A comparison of the estimates of sensitivity between two tests without taking into account the sampling variation of the discrimination limits can increase the type I error to about seven times the nominal value of 0.05. Correct statistical procedures are considered, and the power and required sample size are studied for Gaussian and log-Gaussian distributions of diagnostic test values. The results may be useful for the planning phase of studies to evaluate quantitative diagnostic tests.

Biometry↗

Taking a closer look: time sampling and measurement error.

A person manufactured his in-seat behavior for 15, 30-min sessions so that there were three blocks of five sessions where the behavior occurred 20%, 50%, and 80% of the time. Whole interval, partial interval, and momentary time-sample measures of the behavior were taken and compared to the continuous measure of the behavior i.e., per cent of time the behavior occurred. For interval time sampling, the difference between the continuous and sample measures i.e., measurement error, was: (1) extensive, (2) unidirectional, (3) a function of the time per response, and (4) inconsistent across changes in the continuous measure. A procedural analysis demonstrated that the frequency and duration of behavior are confounded in interval time sampling. Momentary time sampling was found to be superior to interval time sampling in estimating the duration a behavior occurs.

Journal Article↗

Estimation of genotype error rate using samples with pedigree information--an application on the GeneChip Mapping 10K array.

Currently, most analytical methods assume all observed genotypes are correct; however, it is clear that errors may reduce statistical power or bias inference in genetic studies. We propose procedures for estimating error rate in genetic analysis and apply them to study the GeneChip Mapping 10K array, which is a technology that has recently become available and allows researchers to survey over 10,000 SNPs in a single assay. We employed a strategy to estimate the genotype error rate in pedigree data. First, the "dose-response" reference curve between error rate and the observable error number were derived by simulation, conditional on given pedigree structures and genotypes. Second, the error rate was estimated by calibrating the number of observed errors in real data to the reference curve. We evaluated the performance of this method by simulation study and applied it to a data set of 30 pedigrees genotyped using the GeneChip Mapping 10K array. This method performed favorably in all scenarios we surveyed. The dose-response reference curve was monotone and almost linear with a large slope. The method was able to estimate accurately the error rate under various pedigree structures and error models and under heterogeneous error rates. Using this method, we found that the average genotyping error rate of the GeneChip Mapping 10K array was about 0.1%. Our method provides a quick and unbiased solution to address the genotype error rate in pedigree data. It behaves well in a wide range of settings and can be easily applied in other genetic projects. The robust estimation of genotyping error rate allows us to estimate power and sample size and conduct unbiased genetic tests. The GeneChip Mapping 10K array has a low overall error rate, which is consistent with the results obtained from alternative genotyping assays.

Computer Simulation↗

The effect of azone on ocular levobunolol absorption: calculating the area under the curve and its standard error using tissue sampling compartments.

Methods of calculating the area under the concentration-time curve and the associated standard error are proposed for studies in which each animal contributes one independent data point to a pool of data. This approach can be used for data analysis in bioequivalence studies employing tissue sampling compartments. Application of this method indicated that an azone-containing ophthalmic formulation of levobunolol did not produce better ocular bioavailability than a formulation containing no penetration enhancer.

Animals↗

Guidelines for practical utilization of intraoperative frozen sections.

We reviewed 4057 intraoperative frozen sections from 1980 through 1984 to assess the accuracy, strengths, and weaknesses of this technique. Breast, lymph node, and skin comprised half of the sites evaluated. Frozen-section and final diagnoses agreed in 91.5% and disagreed in 6.8% of the cases; 1.7% of the cases were deferred. False-negative frozen-section diagnoses were due to pathologist sampling or judgment errors and surgeon sampling errors. There were eight (0.15%) false-positive diagnoses, none of which altered patient treatment. We recommend that lymph nodes for lymphoproliferative disorders and breast tissue for which a malignant diagnosis will not result in an immediate mastectomy not be submitted for frozen-section diagnosis. Appropriate studies of these tissues can be carried out without an intraoperative diagnosis; such a policy will increase the cost-effectiveness of frozen sections without compromising patient care.

False Negative Reactions↗

Effect of sampling on measurement errors.

Often the analyst is taken as a guarantor for data quality in spite of the fact that sampling is commonly performed by others. If the analyst ignores sampling uncertainties, the money spent on quality control of analysis may sometimes be in vain. The analyst ought to be aware of the difference between controlling exposure and measuring workers' exposure at the workplace. When controlling exposure the aim is to ensure that workers' exposures are below the given occupational exposure limits (OELs); when measuring exposure the aim is to determine what the worker is actually exposed to, on average. In the working environment, exposure is usually controlled by measuring "worst case' situations, i.e., situations where exposure is higher than average by an unknown amount. As pointed out by Eisenhart (cf. Anal. Chem., 1981, 53, 1588A), measuring without a state of statistical control being attained cannot in any logical sense be regarded as measuring anything at all. Except for substances for which the OELs are ceiling limits that must not be exceeded, 'worst case' results cannot be used for documenting non-compliance or for risk assessment, epidemiology or standard setting. Measuring workers' exposure requires estimation of the time weighted average concentration in the exposure period considered (TWAC exposure Period) by carrying out measurements, preferably over a series of days (TWAC Day). Kromhout et al. (Ann. Occup. Hyg., 1993, 37, 253) found TWAC day data to be lognormally distributed with a median geometric standard deviation of 2.5. Sampling from such distributions is shown to give very disperse results. Consequently, many measurement days are needed. A TWAC Exposure Period estimate, therefore, is either very uncertain or has been very costly to obtain. In order to obtain more reliable results at an affordable cost, an alternative approach, called the logbook method, has recently been suggested for the estimation of TWAC Exposure Period. Commonly, workers considered to be similarly exposed are grouped. In contrast, the logbook method groups processes causing similar exposures. The time component of exposure is measured by workers keeping logs of their activities over a period of several weeks.

Air Pollutants, Occupational↗

Maximum likelihood estimates of allele frequencies and error rates from samples of related individuals by gene counting.

SUMMARY: Graphical modeling is used to extend the gene counting method to compute maximum likelihood estimates of allele frequencies for samples of individuals related in extended pedigrees. Genotypes may be missing or partially observed, and error rates can be simultaneously estimated. AVAILABILITY: The Java classes and Javadocs pages for \mathsf\hbox GeneCountAlleles can be obtained from bioinformatics.med.utah.edu/~alun, which also has more information on its use and file formats.

Biological Evolution↗

Increased taxon sampling greatly reduces phylogenetic error.

Several authors have argued recently that extensive taxon sampling has a positive and important effect on the accuracy of phylogenetic estimates. However, other authors have argued that there is little benefit of extensive taxon sampling, and so phylogenetic problems can or should be reduced to a few exemplar taxa as a means of reducing the computational complexity of the phylogenetic analysis. In this paper we examined five aspects of study design that may have led to these different perspectives. First, we considered the measurement of phylogenetic error across a wide range of taxon sample sizes, and conclude that the expected error based on randomly selecting trees (which varies by taxon sample size) must be considered in evaluating error in studies of the effects of taxon sampling. Second, we addressed the scope of the phylogenetic problems defined by different samples of taxa, and argue that phylogenetic scope needs to be considered in evaluating the importance of taxon-sampling strategies. Third, we examined the claim that fast and simple tree searches are as effective as more thorough searches at finding near-optimal trees that minimize error. We show that a more complete search of tree space reduces phylogenetic error, especially as the taxon sample size increases. Fourth, we examined the effects of simple versus complex simulation models on taxonomic sampling studies. Although benefits of taxon sampling are apparent for all models, data generated under more complex models of evolution produce higher overall levels of error and show greater positive effects of increased taxon sampling. Fifth, we asked if different phylogenetic optimality criteria show different effects of taxon sampling. Although we found strong differences in effectiveness of different optimality criteria as a function of taxon sample size, increased taxon sampling improved the results from all the common optimality criteria. Nonetheless, the method that showed the lowest overall performance (minimum evolution) also showed the least improvement from increased taxon sampling. Taking each of these results into account re-enforces the conclusion that increased sampling of taxa is one of the most important ways to increase overall phylogenetic accuracy.

Likelihood Functions↗

An observational study of laterality errors in a sample of clinical records.

BACKGROUND: Confusing left with right eyes can have a potentially serious adverse outcome. The most extreme occurrence is wrong site surgery but even potentially less serious errors can undermine patient confidence in their medical care. This study was designed to look into how often this could be detected in clinical notes. METHODS: An observational study conducted in an ophthalmic hospital. Hundred patients were randomly selected and their clinical notes retrieved. Notes were analysed for the number of left/right transpositions, which part of the notes they were found and whether they were corrected. RESULTS: Forty-four transposition errors were found in 32 sets on notes. The commonest error was drawing the eye on the wrong side of the page. The commonest place where errors were found was in the written outpatient notes. Nineteen of the errors had evidence of later correction. Three consent forms had the incorrect eye denoted and one patient was listed for surgery on the wrong side although this error was corrected before the operation. CONCLUSION: As far as we are aware, this study is the first to look at how often, in standard clinical notes, left/right transposition occurs. Although a direct link cannot made between their occurrence and later wrong side surgery, intuitively it would be reasonable to think it could increase the likelihood if other defences were to fail. We make a number of recommendations that might reduce this confusion and therefore more serious consequences.

Handwriting↗

Sampling and interpretation errors in aerosol monitoring.

The aerosol recorded by simple filter collection or by sophisticated instrument aerosol monitoring may differ considerably from the original, unsampled aerosol. Through use of particle-sizing instruments and computer modeling, the potential biases in sampling, display, and interpretation are demonstrated. The aerosol-size distribution and, therefore, the reported number or mass concentration may be affected by the characteristics of the sampling inlet, the transport to the sensor and the sensor itself. The particle count in specific size ranges determines the precision of the registered particle-size distribution, depending on the weighting chosen. The type of display, by histograms or cumulative plot, focuses on different aspects of the size distribution, and the calibration of the aerosol monitor may modify it further. Particle-size classification to simulate a specific region of the human respiratory system may be achieved through inertial classification or the sensitivity characteristics of the aerosol sensor. Aerosol monitors using passive sampling register the same aerosol-size distribution as active ones, if the aerosol is transported to the sensor with the same efficiency as in the active mode. The sources of various types of errors are presented using computer simulations of typical aerosol-size distributions, often combined with measurements found in the literature. Presentation of these errors in graphical format allows the health professional to estimate more accurately the health implications of aerosol measurements.

Aerosols↗

Comprehensive screening of urine samples for inborn errors of metabolism by electrospray tandem mass spectrometry.

BACKGROUND: Detection of abnormal metabolites in urine is important for the diagnosis of many inborn errors of metabolism (IEM). Rapid, comprehensive screening methods are needed. METHODS: We used electrospray ionization tandem mass spectrometry in positive- and negative-ion modes to detect selected metabolites in urine. For positive-ion analysis, samples were dried and butylated, whereas for negative-ion analysis, samples were merely diluted with the mobile phase. Analysis was by direct injection with multiple reaction monitoring for 32 metabolites in positive mode (amino acids and acylcarnitines) and 30 metabolites in negative mode (organic acids). Run time was 2.1 min in each mode. RESULTS: Interbatch CVs ranged from 4.8% to 32%, enabling quantification of many metabolites. The procedure was applied to controls (278 and 120 in positive- and negative-ion mode, respectively) and 108 IEM individuals representing 37 different IEM. In 105 IEM individuals, representing 36 different IEM, concentrations of one or more diagnostic metabolites were above the 99th percentiles of the control values. CONCLUSIONS: The procedure is faster and less labor-intensive than conventional methods of testing for IEM by amino and organic acid profiling and has similar diagnostic sensitivity. The ability to include a greater range of metabolites offers the potential of a more comprehensive screening procedure.

Amino Acids↗

A method of diagnosing intramammary infection in dairy cows for large experiments.

Diagnosis of microbial infections in the udders of cows in commercial dairy farms for large experiments cannot be without error. Limitations of sampling method and routine prevent collection of the necessary information for sure diagnosis. However, with an organized method of repeated bacteriological examinations using consistent and proven methods of aseptic sampling the errors were shown to be very low. A method based on bacteriological tests on aseptic milk samples was used in 32 herds (approximately 2000 cows) for a 3-year period. This is described and examined in terms of other criteria to validate its use in experimental work. With this method it was not difficult to differentiate between those quarters which regularly shed pathogens and those which did not. Other evidence indicated that it was reasonable to assume that this classification accurately distinguished between infected and uninfected quarters. The errors using this method were quite small: when measuring the state of infection of all quarters in the herds the errors did not exceed 1%. Some small modifications to the method described are suggested to improve further its diagnostic accuracy.

Animals↗

Sampling culturable heterotrophs from microcosms: a statistical analysis.

The contributions of different sources of error in sampling mixed and unmixed bacterial microcosms were evaluated by using analysis of variance. Culturable heterotrophic bacteria from a turbid freshwater impoundment were sampled from 9-liter tanks that were unagitated or mixed with magnetic stirrers or pumps and from dilution bottles that were unagitated or agitated with a mechanical shaker. Axenic cultures of Enterobacter aerogenes were also sampled from manually shaken test tubes. In both agitated and unagitated tanks and in unagitated dilution bottles, dilutions made from the same sampling pipette were significantly different, showing a clumping of bacteria on the scale of millimeters. Also, microcosms within a single experiment differed from one another by a large margin. Dilution mean squares and tank or bottle mean squares were homogeneous for all types of tanks and unagitated bottles, indicating that the gentle mixing provided by pumps and stir bars did not reduce either millimeter scale or intermicrocosm variability over what prevailed in unagitated microcosms. By contrast, the vigorously shaken bottles and test tubes showed no millimeter scale variability. Intermicrocosm variability was undetectable in test tubes and two orders of magnitude less in shaken bottles than in unshaken bottles. When these facts are coupled with the inherent statistical advantage of replicating large rather than small experimental units, it is concluded that sampling error in the enumeration of aquatic bacteria in microcosms will be reduced by using numerous, small, violently agitated microcosms with a minimum of subsampling per microcosm.

Analysis of Variance↗

Corrected small-sample estimation of the Bayes error.

MOTIVATION: A major problem of pattern classification is estimation of the Bayes error when only small samples are available. One way to estimate the Bayes error is to design a classifier based on some classification rule applied to sample data, estimate the error of the designed classifier, and then use this estimate as an estimate of the Bayes error. Relative to the Bayes error, the expected error of the designed classifier is biased high, and this bias can be severe with small samples. RESULTS: This paper provides a correction for the bias by subtracting a term derived from the representation of the estimation error. It does so for Boolean classifiers, these being defined on binary features. Although the general theory applies to any Boolean classifier, a model is introduced to reduce the number of parameters. A key point is that the expected correction is conservative. Properties of the corrected estimate are studied via simulation. The correction applies to binary predictors because they are mathematically identical to Boolean classifiers. In this context the correction is adapted to the coefficient of determination, which has been used to measure nonlinear multivariate relations between genes and design genetic regulatory networks. An application using gene-expression data from a microarray experiment is provided on the website http://gspsnap.tamu.edu/smallsample/ (user:'smallsample', password:'smallsample)').

Algorithms↗