Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sampling Errors”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,243 records · Page 69Linked to original sources

Sampling, log binning, fitting, and plotting durations of open and shut intervals from single channels and the effects of noise.

(1) Analysis of the durations of open and shut intervals measured from single channels currents provides a means to investigate the mechanisms of channel gating. Durations of open and shut intervals are conveniently measured from single channel data by using a threshold level to indicate transitions between open and shut states. This paper presents a detailed characterization of sampling, binning, and noise errors associated with 50% threshold analysis, provides criteria to reduce these errors, methods to correct for them, and presents an efficient means of data handling for binning and plotting interval durations. (2) Measuring interval durations by sampling at a fixed rate introduces two types of errors, (a) the number of intervals of a given measured duration are increased (promoted) over that expected in the absence of sampling, producing a sampling promotion error, (b) sampling decreases the total fraction of true intervals that are detected, producing a sampling detection error. Sampling errors can be reduced to negligible levels if the actual or effective (after interpolation) sampling period is less than 10-20% of both the dead time and fastest time constant in the distribution of intervals. Dead time is given by the duration of a true interval that has a filtered amplitude equal to 50% of the true amplitude. (3) Methods are presented to correct for sampling promotion error during least squares and maximum likelihood fitting. Sampling detection error is more difficult to correct, but an empirical description of the sampling detection error can be used to calculate the effective fraction of detected events with sampling. (4) Noise in the single channel current record can produce two types of error. (a) If noise peaks in the absence of channel activity exceed the threshold for detection, then false channel events of brief duration are produced. Sufficient filtering will prevent this type of error. (b) Noise can also increase the total fraction of true intervals that are detected, producing a noise detection error. Increased filtering over that required to prevent false events is not necessarily the best method for reducing noise detection error, as increased filtering can prevent detection of the faster exponential components. (5) Noise detection error can be reduced in two ways: (a) an empirical description of the noise detection error can be used to calculate the effective fraction of detected events in the presence of noise. (b) The sampling period can be selected so that the sampling detection error cancels the noise detection error.(ABSTRACT TRUNCATED AT 400 WORDS)

Animals↗

An evaluation of cytologic sampling techniques. A comparative study.

False negative cervical cytology is primarily due to errors in sampling. It has been demonstrated that combined ectocervical and endocervical sampling techniques will improve the yield. A prospective study was done to compare ectocervical and combined sampling with a selective technique in which the examiner determines the location of the squamocolumnar junction and chooses the appropriate method. The results demonstrate that combined ectocervical and endocervical sampling significantly increases the number of positive Papanicolaou smears. Selecting which cervices should be samples ectocervically and which need combined sampling does not significantly improve the yield over ectocervical sampling alone. It is the conclusion of this study that both ectocervical and endocervical sampling should routinely be used.

Diagnosis, Differential↗

Discordance between uterine cervical cytology and biopsy: results and etiologies of a one-year audit.

To investigate the etiologies of discrepancies between cervicovaginal smear and corresponding cervical biopsy results, a total of 15,474 cervicovaginal smears were sampled in a one-year period. Among these, 427 patients were diagnosed with atypical squamous cells of undetermined significance (ASCUS), dysplasia, or malignancy. The screen positive rate was 2.8%. All of the positive cases had histologic follow-up. Forty-nine of the 427 patients had a discrepancy of at least two grades (the grades are divided to negative, ASCUS, mild dysplasia, moderate dysplasia, severe dysplasia and invasive carcinoma), between the cytologic and histologic diagnoses. The discrepancy rate was 11.5%. Ten of these discrepant cases had poorly-preserved slides or a not definitely final diagnosis. A total of 39 cases (79.6%) of discrepancy were reviewed in this study. In thirty (77%) of the 39 discrepant cases, the errors were cytologic and in 9 cases (23%) the errors were histologic. Cytologic error was the major cause of cytohistologic discrepancy. The etiologies of cytohistologic discrepancy included: cytologic interpretation error, 17 cases (44%); cytologic sampling error, 10 cases (25%); biopsy sampling error, 6 cases (15%); cytologic screen error, 3 cases (8%); and biopsy interpretation error, 3 cases (8%). The major etiology of cytohistologic discordances was cytologic interpretation error. In this retrospective study, we determined the etiologies of cytohistologic discrepancies. This information can be useful for improving diagnostic accuracy and the quality of patient care.

Biopsy↗

Data quality objectives for surface-soil cleanup operation using in situ gamma spectrometry for concentration measurements.

In situ gamma spectrometry is an efficient method for monitoring the progress of cleanup activities for radioactive contaminants in surface soil and for evaluating the attainment of cleanup standards. However, desired data precision and accuracy must be specified for such a detection system prior to the operation to ensure that the level of uncertainty associated with the concentration measurements is acceptable. A method for developing data quality objectives is described in this paper for in situ gamma spectrometry to achieve numerical goals for data precision and accuracy for cleanup operations. Concentration measurement for a radionuclide at its cleanup level must have a precision commensurate with the importance of cleanup decisions. The 95% lower limit of detection of the system is suggested to be about one tenth the expected system response at the cleanup level. The count time required to achieve the preferred 95% lower limit of detection, and hence the desired precision, can then be determined. The accuracy error arises from the overall calibration factor, which relates the detector responses (e.g., count rate) to physical quantities of interest (e.g., radionuclide soil concentration). The major source of error for the calibration factor using in situ gamma spectrometry is the misidentification of the type of the depth profile of radionuclide concentration in soil. If surrogate radionuclides are used, such as 241Am for plutonium, the variation in the concentration ratio would be another significant source of error. Soil sampling programs performed prior to a cleanup operation will greatly reduce the accuracy error for an in situ detection system, and the analysis of system errors may determine the degree of sampling required. The planning of such a program is discussed in the study. Uncertainty analysis using a Latin Hypercube sampling technique for the calibration factor is also demonstrated. The quantitative result of the uncertainty analysis is useful for determining a nuclide's maximum peak count rate using gamma spectrum that ensures the attainment of the cleanup standard for that nuclide with a pre-specified confidence level (e.g., 95%). The cleanup operation of 239,240Pu in surface soil in the safety shot areas at the Nevada Test Site serves as an example to illustrate the data quality objectives development.

Americium↗

Near-infrared analysis of fat, protein, and casein in cow's milk.

Fat, crude protein, true protein, and casein were determined in cow milks by near-infrared transmission spectroscopy (NIR). Partial and overall PLS calibrations were performed on two sets of samples: partial calibration included 76 unhomogenized samples, whereas overall calibration used 96 homogenized and unhomogenized samples. Standard errors of calibration were 0.12% for fat, 0.06% for crude protein, 0.04% for true protein, and 0.05% for casein in the overall calibration. Validation of the overall calibration with an independent set of samples gave standard errors of prediction of 0. 07% for fat, 0.06% for crude protein and casein, and 0.05% for true protein. Except for fat, all of the statistical parameters were better with overall than with partial calibrations, which indicates that homogenization has an effect on NIR fat determination. Despite the relatively small number of samples included in the calibration model, NIR transmission was found to be a reliable method for the determination of fat and nitrogenous constituents in milk.

Animals↗

Effect of analytical run length on quality-control (QC) performance and the QC planning process.

The performance measure traditionally used in the quality-control (QC) planning process is the probability of rejecting an analytical run when an out-of-control error condition exists. A shortcoming of this performance measure is that it doesn't allow comparison of QC strategies that define analytical runs differently. Accommodating different analytical run definitions is straightforward if QC performance is measured in terms of the average number of patient samples to error detection, or the average number of patient samples containing an analytical error that exceeds total allowable error. By using these performance measures to investigate the impact of different analytical run definitions on QC performance demonstrates that during routine QC monitoring, the length of the interval between QC tests can have a major influence on the expected number of unacceptable results produced during the existence of an out-of-control error condition.

Chemistry, Clinical↗

Probabilistic estimation of microarray data reliability and underlying gene expression.

BACKGROUND: The availability of high throughput methods for measurement of mRNA concentrations makes the reliability of conclusions drawn from the data and global quality control of samples and hybridization important issues. We address these issues by an information theoretic approach, applied to discretized expression values in replicated gene expression data. RESULTS: Our approach yields a quantitative measure of two important parameter classes: First, the probability P(sigma|S) that a gene is in the biological state sigma in a certain variety, given its observed expression S in the samples of that variety. Second, sample specific error probabilities which serve as consistency indicators of the measured samples of each variety. The method and its limitations are tested on gene expression data for developing murine B-cells and a t-test is used as reference. On a set of known genes it performs better than the t-test despite the crude discretization into only two expression levels. The consistency indicators, i.e. the error probabilities, correlate well with variations in the biological material and thus prove efficient. CONCLUSIONS: The proposed method is effective in determining differential gene expression and sample reliability in replicated microarray data. Already at two discrete expression levels in each sample, it gives a good explanation of the data and is comparable to standard techniques.

Algorithms↗

Power and sample size calculations in the presence of phenotype errors for case/control genetic association studies.

BACKGROUND: Phenotype error causes reduction in power to detect genetic association. We present a quantification of phenotype error, also known as diagnostic error, on power and sample size calculations for case-control genetic association studies between a marker locus and a disease phenotype. We consider the classic Pearson chi-square test for independence as our test of genetic association. To determine asymptotic power analytically, we compute the distribution's non-centrality parameter, which is a function of the case and control sample sizes, genotype frequencies, disease prevalence, and phenotype misclassification probabilities. We derive the non-centrality parameter in the presence of phenotype errors and equivalent formulas for misclassification cost (the percentage increase in minimum sample size needed to maintain constant asymptotic power at a fixed significance level for each percentage increase in a given misclassification parameter). We use a linear Taylor Series approximation for the cost of phenotype misclassification to determine lower bounds for the relative costs of misclassifying a true affected (respectively, unaffected) as a control (respectively, case). Power is verified by computer simulation. RESULTS: Our major findings are that: (i) the median absolute difference between analytic power with our method and simulation power was 0.001 and the absolute difference was no larger than 0.011; (ii) as the disease prevalence approaches 0, the cost of misclassifying a unaffected as a case becomes infinitely large while the cost of misclassifying an affected as a control approaches 0. CONCLUSION: Our work enables researchers to specifically quantify power loss and minimum sample size requirements in the presence of phenotype errors, thereby allowing for more realistic study design. For most diseases of current interest, verifying that cases are correctly classified is of paramount importance.

Alzheimer Disease↗

Misuse of standard error of the mean (SEM) when reporting variability of a sample. A critical evaluation of four anaesthesia journals.

BACKGROUND: In biomedical research papers, authors often use descriptive statistics to describe the study sample. The standard deviation (SD) describes the variability between individuals in a sample; the standard error of the mean (SEM) describes the uncertainty of how the sample mean represents the population mean. Authors often, inappropriately, report the SEM when describing the sample. As the SEM is always less than the SD, it misleads the reader into underestimating the variability between individuals within the study sample. METHODS: The aim of this study was to evaluate the frequency of inappropriate use of the SEM in four leading anaesthesia journals in 2001. The journals were searched manually for descriptive statistics reporting either the mean (SD) or the mean (SEM), and inappropriate use of the SEM was noted. RESULTS: In 2001, all four anaesthesia journals published articles that used the SEM incorrectly: Anesthesia & Analgesia 27.7%, British Journal of Anaesthesia 22.6%, Anesthesiology 18.7% and European Journal of Anaesthesiology 11.5%. Laboratory reports and clinical studies were equally affected, except for Anesthesiology where 90% were basic science reports. CONCLUSIONS: One in four articles (n=198/860, 23%) published in four anaesthesia journals in 2001 inappropriately used the SEM in descriptive statistics to describe the variability of the study sample. Anaesthesia journals are encouraged to provide clearer statistical guidelines on how to report data variability in descriptive statistics.

Anesthesiology↗

Hypothesis testing in neurosurgical trials.

Controlled clinical trials represent the most scientific methods of evaluating a new form of treatment. In designing such a trial, one must avoid committing two kinds of errors. The Type I error is defined as falsely concluding that a difference between two treatments exists, when they are equal. The Type II error is committed when one concludes that two treatments are the same, when a real difference exists. To reduce the probability of committing these errors, large sample sizes are required. A survey of neurosurgical trials showed that the majority of these trials have an unacceptably high probability of committing a Type II error because of inadequate sample size.

Clinical Trials as Topic↗

Correlation between fetal scalp blood samples and intravascular blood pH, pO2 and oxygen saturation measurements.

OBJECTIVE: This study was designed to compare the values of blood gases in local scalp blood, obtained by scalp blood sampling, with direct measurements of fetal preductal arterial blood in fetal sheep. METHODS: Six fetal lambs were catheterized in the brachial artery and vein. Maternal oxygenation was reduced in steps from a fractional inspired oxygen concentration (FiO2) of 20 to 5% by addition of nitrogen to the inhaled gas mixture. Fetal scalp and arterial blood were sampled simultaneously at maternal FiO2 step intervals after maternal FiO2 and oxygen were stable for > 5 min. Blood pH, pO2 and oxygen saturation were measured and linear regression was performed to determine the correlation between these values. RESULTS: Scalp pH correlated well with arterial pH, whereas scalp pO2, pCO2 and oxygen saturation did not. However, when a secondary analysis was performed taking into account the effects of aerobic contamination, scalp pCO2, pO2 and oxygen saturation became highly correlated with arterial values. CONCLUSIONS: Local scalp blood oxygen saturation correlates highly with fetal preductal arterial values, when physiological artifacts are eliminated. The technique of scalp blood sampling introduces error into oxygenation saturation measurements, owing to difficulties in anaerobic sample collection. These results suggest that continuous measurement of fetal scalp oxygenation by noninvasive oximetry may be superior to direct sampling of scalp blood.

Animals↗

Temporal frequency analysis of dynamic MRI techniques.

Dynamic imaging strategies often involve updating certain areas of k-space (i.e., the low spatial frequencies) more frequently than others. However, important dynamic signal changes may occur anywhere in k-space. In this study, a dynamic k-space sampling analysis method was developed to determine the energy error associated with specific dynamic sampling strategies. The method uses the temporal power spectrum of k-space signals to determine the level and k-space locations of sampling errors. The proposed method was used to compare two dynamic sampling strategies (full sequential and keyhole) for a dynamic first-pass bolus simulation and a continuous heart imaging study. The error analysis agreed well with the errors in the reconstructed images. The technique can be used to determine the minimum sampling frequency for any location in the k-space, and may ultimately be used to optimize dynamic sampling strategies. Magn Reson Med 45:550-556, 2001.

Computer Simulation↗

A multidisciplinary approach to enhance documentation of antibiotic serum sampling.

A procedure to improve interdepartmental communication and documentation of antibiotic serum sampling data for pharmacokinetic evaluation will be presented. A prospective audit by the Pharmacokinetic Service revealed that approximately 40% of all antibiotic serum levels were improperly drawn resulting in unsuitable specimens and erroneous serum concentrations or lacked sufficient data for pharmacokinetic analysis. A lack of communication and documentation between phlebotomy and nursing personnel was found to be the most significant source of potential error in serum sampling. Once the protocol for serum sampling was revised, less than 5% of antibiotic serum levels were found to be unsuitable for evaluation and interpretation. A continuous audit for procedural compliance identifies any source of potential sampling error and provides a means to improve the overall quality of a Pharmacokinetic Service.

Anti-Bacterial Agents↗

DNA barcoding: error rates based on comprehensive sampling.

DNA barcoding has attracted attention with promises to aid in species identification and discovery; however, few well-sampled datasets are available to test its performance. We provide the first examination of barcoding performance in a comprehensively sampled, diverse group (cypraeid marine gastropods, or cowries). We utilize previous methods for testing performance and employ a novel phylogenetic approach to calculate intraspecific variation and interspecific divergence. Error rates are estimated for (1) identifying samples against a well-characterized phylogeny, and (2) assisting in species discovery for partially known groups. We find that the lowest overall error for species identification is 4%. In contrast, barcoding performs poorly in incompletely sampled groups. Here, species delineation relies on the use of thresholds, set to differentiate between intraspecific variation and interspecific divergence. Whereas proponents envision a "barcoding gap" between the two, we find substantial overlap, leading to minimal error rates of approximately 17% in cowries. Moreover, error rates double if only traditionally recognized species are analyzed. Thus, DNA barcoding holds promise for identification in taxonomically well-understood and thoroughly sampled clades. However, the use of thresholds does not bode well for delineating closely related species in taxonomically understudied groups. The promise of barcoding will be realized only if based on solid taxonomic foundations.

Animals↗

Quantifying the percent increase in minimum sample size for SNP genotyping errors in genetic model-based association studies.

Kang et al. [Genet Epidemiol 2004;26:132-141] addressed the question of which genotype misclassification errors are most costly, in terms of minimum percentage increase in sample size necessary (%MSSN) to maintain constant asymptotic power and significance level, when performing case/control studies of genetic association in a genetic model-free setting. They answered the question for single nucleotide polymorphisms (SNPs) using the 2 x 3 chi2 test of independence. We address the same question here for a genetic model-based framework. The genetic model parameters considered are: disease model (dominant, recessive), genotypic relative risk, SNP (marker) and disease allele frequency, and linkage disequilibrium. %MSSN coefficients of each of the six possible error rates are determined by expanding the non-centrality parameter of the asymptotic distribution of the 2 x 3 chi2 test under a specified alternative hypothesis to approximate %MSSN using a linear Taylor series in the error rates. In this work we assume errors misclassifying one homozygote as another homozygote are 0, since these errors are thought to rarely occur in practice. Our findings are that there are settings of the genetic model parameters that lead to large total %MSSN for both dominant and recessive models. As SNP minor allele approaches 0, total %MSSN increases without bound, independent of other genetic model parameters. In general, %MSSN is a complex function of the genetic model parameters. Use of SNPs with small minor allele frequency requires careful attention to frequency of genotyping errors to insure that power specifications are met. Software to perform these calculations for study design is available, and an example of its use to study a disease is given.

Alleles↗

Cervical cancers diagnosed after negative results on cervical cytology: perspective in the 1980s.

OBJECTIVES: To assess the magnitude of the problem of interval cancers of the cervix (those that are diagnosed within a short time after negative screening test results) in the 1980s, to compare the nature of interval cancers in younger women with that in older women, and, by reviewing negative cervical smears, to determine the proportion of interval cancers that might represent the development of malignancy anew compared with the proportion that might be associated with difficulties in sampling or errors in reporting. DESIGN: An audit of the interval cases of cervical cancer that had been diagnosed within 36 months of a smear having been reported as negative by the Victorian Cytology Gynaecological Service among women registered with cervical cancer during 1982-6. SETTING: The Victorian Cytology Gynaecological Service, a free public sector cytology laboratory in Victoria, Australia. SUBJECTS: 138 Women, all of whom had had cervical cancer diagnosed during the 36 months after having had a negative cervical smear. Subjects were divided into two age groups: younger women, aged less than 35; older women, aged 35-69. INTERVENTIONS: Negative slides were reviewed for evidence of optimal sampling and for the presence of cellular abnormalities that had been missed at the time of the original reporting. MAIN OUTCOME MEASURES: The number of interval cases of cancer of the cervix registered during 1982-6. The proportion of interval cases occurring in younger women and the proportion occurring in older women. Division of women into three risk categories based on clinical history and screening history that broadly corresponded to the probability that a diagnosis of cervical cancer might be expected during the 36 months after the issuing of a negative smear report. RESULTS: 138 Of 1044 (13.2%) women who had been registered with cervical cancer during 1982-6 had had one or more negative smears during the 36 months preceding the diagnosis of cancer. Interval cancers comprised a larger proportion of registrations of cervical cancer in women aged less than 35 years than in women aged 35-69 (21.1% v 11.0%, p less than 0.01). Women with interval cancer who had had at least three negative smears during the 10 years before the diagnosis of cancer were commoner in the younger age group than in the older age group (7.0% v 2.5%, p less than 0.01). When, however, the number of observed cases of squamous cell carcinoma was related to the number of expected cases in the absence of screening, no significant difference was found between the two age groups (6.8% v 4.8%, p greater than 0.10). The rate of diagnosis of interval cancer per 100,000 negative tests was lower among younger women than among older women (10/100,000 v 16/100,000). Review of the negative slides showed that 11.9% were again considered to be negative with an optimal sample having been obtained as evidenced by the presence of endocervical cells or metaplastic cells, or both. CONCLUSIONS: Interval cancers might comprise a larger proportion of all registered cases of cervical cancer among younger women owing to the larger proportion of such cancers being prevented in this age group. Among women with interval cancer review of the negative slides showed that most were accounted for by suboptimal sampling or by errors of reporting.

Adult↗

Statistical methods for describing occupational exposure measurements.

An important step in studies relating worker health to industrial exposure is the estimation of mean exposure levels. The investigator frequently has to rely on industrial hygiene measurements collected for other purposes. Samples may have been taken at several companies on different dates, and on each occasion multiple individual samplers may have been employed. Often it is not recognized that readings from such a hierarchical arrangement are correlated; for example, samples taken at the same time and location are more alike than samples taken on different days. This correlation invalidates the commonly used standard errors of sample means and the usual sample standard deviation. A component of variance analysis is suggested which quantifies within-day, between-day and between-company variation. Estimators of mean exposure are presented with correct standard errors. The techniques are illustrated by a small set of data and by a recent study of exposures to styrene in 36 companies manufacturing reinforced plastics.

Environmental Exposure↗