Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sampling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,729 records · Page 96Linked to original sources

Determination of closed form solution for acceptance sampling using ANN.

Tabled sampling schemes such as MIL-STD-105D offer limited flexibility to quality control engineers in designing sampling plans to meet specific needs. We describe a closed form solution to determine the AQL indexed single sampling plan using an artificial neural network (ANN). To determine the sample size and the acceptance number, feed-forward neural networks with sigmoid neural function are trained by a back propagation algorithm for normal, tightened, and reduced inspections. From these trained ANNs, the relevant weight and bias values are obtained. The closed form solutions to determine the sampling plans are obtained using these values. Numerical examples are provided for using these closed form solutions to determine sampling plans for normal, tightened, and reduced inspections. The proposed method does not involve table look-ups or complex calculations. Sampling plan can be determined by using this method, for any required acceptable quality level and lot size. Suggestions are provided to duplicate this idea for applying to other standard sampling table schemes.

Algorithms↗

Sample preparation followed by high performance liquid chromatographic (HPLC) analysis for monitoring muconic acid as a biomarker of occupational exposure to benzene.

Factors affecting solid phase extraction (SPE) of trans,trans-muconic acid (ttMA), as a benzene biomarker, including sample pH, sample concentration, sample volume, sample flow rate, washing solvent, elution solvent, and type of sorbent were evaluated. Extracted samples were determined by HPLC-UV (high performance liquid chromatography-ultraviolet). The analytical column was C18, UV wave length was 259 nm, and the mobile phase was H(2)O/methanol/acetic acid run at flow rate of 1 ml/min. A strong anion exchange silica cartridge was found successful in simplifying SPE. There was a significant difference between recoveries of ttMA when different factors were used (p < .001). An optimum recovery was obtained when sample pH was adjusted at 7. There was no significant difference when different sample concentrations were used (p > .05). The optimized method was then validated with 3 different pools of samples showing good reproducibility over 6 consecutive days and 6 within-day experiments.

Air Pollutants, Occupational↗

A task-based statistical model of a worker's exposure distribution: Part II--Application to sampling strategy.

A task-based statistical model of a worker's exposure distribution for an airborne chemical toxicant is applied to estimating the long-term average exposure level, mu. The precision in estimation is represented by the variance of the sample estimator, denoted by Var[mu]. A traditional sampling strategy consists of integratively measuring the 8-hr time-weighted average exposure level on randomly selected workdays, and computing the sample mean; this strategy is termed "simple one-stage cluster sampling," where each 8-hr workday is a cluster of thirty-two 15-min periods. Three alternative strategies involving measurements of 15-min TWAs are examined: simple random sampling of 15-min periods, and stratified random sampling of 15-min periods with proportional allocation by task, and with optimum allocation by task. All four survey designs provide unbiased estimates of mu. However, for a fixed cost, the stratified sampling designs may provide a lower Var[mu] than simple one-stage cluster sampling for less work time monitored.

Air Pollutants, Occupational↗

Development of a sampling and analytical method for 2,2-dichloro-1,1,1-trifluoroethane in workplace air.

The use of 2,2-dichloro-1,1,1-trifluoroethane (compound number: HCFC-123) is growing in industry as a substitute for ozone-depleting chlorofluorocarbons (CFCs). Recently, liver-related illnesses have been reported from industries handling HCFC-123. However, information on worker exposure to the material is limited, and an acceptable sampling/analytical method is not available. The aim of this study was to develop a widely applicable sampling and analytical method to determine worker exposures to airborne HCFC-123 and to evaluate the performance of the method. A solid sorbent tube, containing two sections (400 mg in the front and 200 mg in the back) of activated coconut-shell charcoal was chosen for sampling airborne HCFC-123 vapor. The breakthrough volumes were 13.6 L at 3597 +/- 210.1 ppm (with a sampling airflow rate of 0.046 L/min) and 17.0 L at 1841 +/- 4.5 ppm (with sampling airflow rate of 0.046-0.050 L/min). Samples of HCFC-123 in the charcoal tube were stable for 7 days either at room temperature or in a refrigerator and a migration occurred within 14 days at room temperature. It is recommended that the HCFC-123 sample in activated charcoal tubes be stored either at room temperature or in a refrigerator and be analyzed within 7 days. The HCFC-123 in the charcoal tubes was desorbed into dichloromethane and analyzed using gas chromatography/ flame ionization detection. The limit of detection was 0.23 mg/sample, and the average desorption efficiency was 99.0%. The total coefficient of variation was 0.060, and the method accuracy was 16.6%. In conclusion, the performance of the sampling and analytical method developed for the determination of airborne HCFC-123 concentrations was acceptable to the NIOSH sampling and analytical criteria.

Air Pollutants, Occupational↗

Strong feature sets from small samples.

For small samples, classifier design algorithms typically suffer from overfitting. Given a set of features, a classifier must be designed and its error estimated. For small samples, an error estimator may be unbiased but, owing to a large variance, often give very optimistic estimates. This paper proposes mitigating the small-sample problem by designing classifiers from a probability distribution resulting from spreading the mass of the sample points to make classification more difficult, while maintaining sample geometry. The algorithm is parameterized by the variance of the spreading distribution. By increasing the spread, the algorithm finds gene sets whose classification accuracy remains strong relative to greater spreading of the sample. The error gives a measure of the strength of the feature set as a function of the spread. The algorithm yields feature sets that can distinguish the two classes, not only for the sample data, but for distributions spread beyond the sample data. For linear classifiers, the topic of the present paper, the classifiers are derived analytically from the model, thereby providing an enormous savings in computation time. The algorithm is applied to cancer classification via cDNA microarrays. In particular, the genes BRCA1 and BRCA2 are associated with a hereditary disposition to breast cancer, and the algorithm is used to find gene sets whose expressions can be used to classify BRCA1 and BRCA2 tumors.

Breast Neoplasms↗

Methodological rigor with internet samples: new ways to reach underrepresented populations.

We present several rigorous methods for sampling difficult-to-reach and empirically underrepresented populations via the Internet. The methodology's representativeness was tested by comparing the demographics of a small sample of 82 lesbian and bisexual females with a much larger Gallup Organization sample of the general population (n > 1,000) obtained via random digit dialing. Compared to the latter poll, the rigorous sampling designs developed for the Internet were found to be significantly more robust and equally representative of the U.S. general population. The Gallup Organization reached a sample more representative of the age distribution of the United States. The Internet sample reached a sample more representative of the population, with less education, lower incomes, and a broad spectrum of ethnic diversity. The samples were equally effective in representing the distribution of the population with rural and urban residence.

Adolescent↗

Sample size for identifying differentially expressed genes in microarray experiments.

Microarray technology allows simultaneous comparison of expression levels of thousands of genes under each condition. This paper concerns sample size calculation in the identification of differentially expressed genes between a control and a treated sample. In a typical experiment, only a fraction of genes (altered genes) is expected to be differentially expressed between two samples. Sample size determination depends on a number of factors including the specified significance level (alpha), the desired statistical power (1-beta), the fraction (eta) of truly altered genes out of the total g genes studied, and the effect sizes (Delta) for the altered genes. This paper proposes a method to calculate the number of arrays required to detect at least 100lambda % (where 0 < lambda < or = 1) of the truly altered genes under the model of an equal effect size for all altered genes. The required numbers of arrays are tabulated for various values of alpha, beta, Delta, eta, and lambda for the one-sample and two-sample t-tests for g = 10,000. Based on the proposed approach, to identify up to 90% of truly altered genes among the unknown number of truly altered genes, the estimated numbers of arrays needed appear to be manageable. For instance, when the standardized effect size is at least 2.0, the number of arrays needed is less than or equal to 14 for the two-sample t-test and is less than or equal to 10 for the one-sample t-test. As the cost per array declines, such array numbers become practical. The proposed method offers a simple, intuitive, and practical way to determine the number of arrays needed in microarray experiments in which the true correlation structure among the genes under investigation cannot be reasonably assumed. An example dataset is used to illustrate the use of the proposed approach to plan microarray experiments.

Animals↗

Optimal number of features as a function of sample size for various classification rules.

MOTIVATION: Given the joint feature-label distribution, increasing the number of features always results in decreased classification error; however, this is not the case when a classifier is designed via a classification rule from sample data. Typically (but not always), for fixed sample size, the error of a designed classifier decreases and then increases as the number of features grows. The potential downside of using too many features is most critical for small samples, which are commonplace for gene-expression-based classifiers for phenotype discrimination. For fixed sample size and feature-label distribution, the issue is to find an optimal number of features. RESULTS: Since only in rare cases is there a known distribution of the error as a function of the number of features and sample size, this study employs simulation for various feature-label distributions and classification rules, and across a wide range of sample and feature-set sizes. To achieve the desired end, finding the optimal number of features as a function of sample size, it employs massively parallel computation. Seven classifiers are treated: 3-nearest-neighbor, Gaussian kernel, linear support vector machine, polynomial support vector machine, perceptron, regular histogram and linear discriminant analysis. Three Gaussian-based models are considered: linear, nonlinear and bimodal. In addition, real patient data from a large breast-cancer study is considered. To mitigate the combinatorial search for finding optimal feature sets, and to model the situation in which subsets of genes are co-regulated and correlation is internal to these subsets, we assume that the covariance matrix of the features is blocked, with each block corresponding to a group of correlated features. Altogether there are a large number of error surfaces for the many cases. These are provided in full on a companion website, which is meant to serve as resource for those working with small-sample classification. AVAILABILITY: For the companion website, please visit http://public.tgen.org/tamu/ofs/ CONTACT: e-dougherty@ee.tamu.edu.

Algorithms↗

Effect of pooling samples on the efficiency of comparative studies using microarrays.

MOTIVATION: Many biomedical experiments are carried out by pooling individual biological samples. However, pooling samples can potentially hide biological variance and give false confidence concerning the data significance. In the context of microarray experiments for detecting differentially expressed genes, recent publications have addressed the problem of the efficiency of sample pooling, and some approximate formulas were provided for the power and sample size calculations. It is desirable to have exact formulas for these calculations and have the approximate results checked against the exact ones. We show that the difference between the approximate and the exact results can be large. RESULTS: In this study, we have characterized quantitatively the effect of pooling samples on the efficiency of microarray experiments for the detection of differential gene expression between two classes. We present exact formulas for calculating the power of microarray experimental designs involving sample pooling and technical replications. The formulas can be used to determine the total number of arrays and biological subjects required in an experiment to achieve the desired power at a given significance level. The conditions under which pooled design becomes preferable to non-pooled design can then be derived given the unit cost associated with a microarray and that with a biological subject. This paper thus serves to provide guidance on sample pooling and cost-effectiveness. The formulation in this paper is outlined in the context of performing microarray comparative studies, but its applicability is not limited to microarray experiments. It is also applicable to a wide range of biomedical comparative studies where sample pooling may be involved.

Computational Biology↗

Optimal sampling strategies for early pharmacodynamic measures in tuberculosis.

OBJECTIVES: To evaluate whether methodological optimization of serial sputum colony counting (SSCC) studies, a potentially important component in the drug development process for tuberculosis, could significantly improve their power. METHODS: Simulations were carried out using a model derived from a large SSCC dataset. Variance inflation factors (VIFs) were calculated for model parameters, focusing on the elimination rate constant likely to reflect 'sterilizing' activity and sampling schemes were optimized relative to a scheme of daily sampling during the initial phase of therapy. Corresponding sample sizes required for SSCC studies using different schemes were also computed. RESULTS: Published sampling schemes lacked efficiency with respect to the 'sterilizing' phase. Pragmatic optimized schemes yielding greatest precision were achieved using eleven sampling points around a skeleton of 0, 2, 7, 14 and 56 days. The standard error of the 'sterilizing' rate constant was reduced more than 4-fold, and sample size for realistic treatment effects was effectively halved. Even schemes with a restricted duration of sampling to avoid high proportions of missing data and those with fewer sampling points still achieved significant gains in precision. Sensitivity analysis suggested that such schemes should continue to perform well over the immediately foreseeable range of improvements in therapy. CONCLUSIONS: Methodological improvements in the design of SSCC studies could make them a powerful tool in Phase II development of anti-tuberculosis agents.

Antitubercular Agents↗

Sampling issues in research with nonorganic failure-to-thrive children.

Describes the methodological problems posed by sampling characteristics of nonorganic failure-to-thrive (NOFT) children and strategies to address them. Sources of variance in sample characteristics include the criteria used to define NOFT, populations from which samples are drawn, parent refusal and sample attrition, medical and psychological treatment, and individual differences in environmental or biologic risk factors. Undesirable consequences of unrecognized intrasample variation in NOFT include sample bias, limited generalizability of findings across different settings, and erroneous assumption of sample homogeneity. Comprehensive description of sample selection and characteristics will enhance more accurate definition of NOFT, generalizability of findings, and evaluation of the impact of sampling characteristics. Future studies should focus on objective assessment of subtypes of NOFT and the relationship of individual difference variables to psychological status and prognosis.

Diagnosis, Differential↗

Gene genealogies when the sample size exceeds the effective size of the population.

We study the properties of gene genealogies for large samples using a continuous approximation introduced by R. A. Fisher. We show that the major effect of large sample size, relative to the effective size of the population, is to increase the proportion of polymorphisms at which the mutant type is found in a single copy in the sample. We derive analytical expressions for the expected number of these singleton polymorphisms and for the total number of polymorphic, or segregating, sites that are valid even when the sample size is much greater than the effective size of the population. We use simulations to assess the accuracy of these predictions and to investigate other aspects of large-sample genealogies. Lastly, we apply our results to some data from Pacific oysters sampled from British Columbia. This illustrates that, when large samples are available, it is possible to estimate the mutation rate and the effective population size separately, in contrast to the case of small samples in which only the product of the mutation rate and the effective population size can be estimated.

DNA, Mitochondrial↗

Qualitative and quantitative assessment of geographic clustering of population samples selected using different methods of random digit dialing.

Random digit dialing is a method commonly used to select random population samples in epidemiologic research. Although random digit dialing is generally presumed to provide representative samples, coverage and nonresponse errors may affect the degree to which the sample is representative. The present investigation was undertaken to determine whether the geographic distributions of samples selected using variations of the basic random digit dialing technique accurately reflect the underlying population distribution, and if not, whether such samples tend to be either more or less dispersed than the populations from which they were selected. Data regarding control groups from three case-control studies conducted from 1983-1986 were utilized. The residence addresses of 998 controls were located and assigned an X-Y coordinate and census tract designation within a three-county geographic area in northwest Washington State. Initially, the spatial distributions of controls were examined graphically in relation to the age-sex structure of the underlying population. No differences in geographic pattern were apparent. A more formal statistical evaluation was conducted based on centrographic techniques utilizing two measures to describe the spatial distribution of a set of points. Results indicate that the samples chosen were neither more nor less dispersed than the underlying populations. However, the geographic centers of samples selected using primary numbers tended to be shifted from the centers of their respective populations. Several possible explanations for such shifts are considered, and extensions of the analytical approach are suggested in relation to the further evaluation of population sampling and the investigation of space-time aggregations of disease.

Adult↗

A methodology for sampling and accessing homeless individuals in Melbourne, 1995-96.

A methodology for sampling homeless populations in inner Melbourne was developed to study their health status and prevalence of tuberculosis. This paper describes the design, development and implementation of the project. The results of health status and tuberculosis analysis are published elsewhere. Involvement and interaction with local service providers and agencies to homeless people was central to the project throughout. A definitional construct of homelessness was developed, drawn from local and overseas literature and contemporary local experience. The study's aim was to obtain a representative sample of homeless individuals in various levels of accommodation and a convenience sample of those who were unaccommodated (streets and parks). A comprehensive sampling frame of accommodation options was constructed from available databases, and systematic sampling applied to produce a sample of 396 beds, from which 284 participants were enrolled. Convenience sampling of unaccommodated homeless individuals produced 100 participants. All agreed to undergo a comprehensive questionnaire, blood and Mantoux testing, the latter being completed successfully in 94%. Commonsense, cultural sensitivity and a non-threatening approach were critical to the success of the project and the security of the field workers. The methods described attempt to address recognised difficulties of sampling from homeless populations and should be reproducible both in the future and elsewhere. Potential for selection bias remains the main threat to validity, which the described methodology combined with adequate resources should help to address.

Health Surveys↗

Complex sampling: implications for data analysis.

Investigators in dental public health often use strategies other than simple random sampling to identify potential subjects; however, their statistical analyses do not always take into account the complex sampling mechanism. Often it is not clear whether a given strategy requires adjustment for stratification and/or cluster sampling of observations. We propose that the need for such adjustment depends on the primary study objective. As a general rule, we recommend that if the study goal is to estimate the magnitude of either a population value of interest (e.g., prevalence), or an established exposure-outcome association, adjustment of variances to reflect complex sampling is essential because obtaining appropriate variance estimates is a priority. However, if the study goal is to establish the presence of an association, especially in a preliminary investigation of novel conditions or understudied populations, obtaining appropriate variance estimates may not be of primary importance; hence, adjustment of variances for complex sampling is not always required, but often is recommended. This paper describes several types of complex sampling designs, methods of adjusting for complex sampling strategies, examples illustrating the effect of adjustment, and alternative approaches for analysis of complex samples.

Analysis of Variance↗

Optimal sampling schedule design for populations of patients.

Generation of pharmacodynamic relationships in the clinical arena requires estimation of pharmacokinetic parameter values for individual patients. When the target population is severely ill, the ability to obtain traditional intensive blood sampling schedules is curtailed. Population modeling guided by optimal sampling theory has provided robust estimates of individual patient pharmacokinetic parameter values. Because of the wide range of parameter values seen in this circumstance, it is important to know how the range of parameter values in the population affects the timing of the optimal samples. We describe a new, simple technique to obtain optimal samples for a population of patients. This technique uses the nonparametric distribution associated with a nonparametric adaptive grid population pharmacokinetic analysis. We used the distribution from an analysis of 58 patients receiving levofloxacin for nosocomial pneumonia at a dose of 750 mg. The collection of parameter vectors and their associated probabilities were entered into a D-optimal design evaluation by using ADAPT II. The sampling times, weighted for their probabilities, were displayed in a frequency histogram (an expression of how system information varies with time for the population). Such an explicit expression of the time distribution of information allows rational sampling design that is robust not only for the population mean vector, as in traditional D-optimal design theory, but also for large portions of the total population. For levofloxacin, one reasonable six-sample design would be 1.5, 2, 2.25, 4, 4.75, and 24 h after starting a 90-min infusion. Such sampling designs allow informative population pharmacokinetic analysis with precise and unbiased estimates after the maximal a posteriori probability Bayesian step. This allows the highest probability of delineating a pharmacodynamic relationship.

Chromatography, High Pressure Liquid↗

Sample sizes of studies on diagnostic accuracy: literature survey.

OBJECTIVES: To determine sample sizes in studies on diagnostic accuracy and the proportion of studies that report calculations of sample size. DESIGN: Literature survey. DATA SOURCES: All issues of eight leading journals published in 2002. METHODS: Sample sizes, number of subgroup analyses, and how often studies reported calculations of sample size were extracted. RESULTS: 43 of 8999 articles were non-screening studies on diagnostic accuracy. The median sample size was 118 (interquartile range 71-350) and the median prevalence of the target condition was 43% (27-61%). The median number of patients with the target condition--needed to calculate a test's sensitivity--was 49 (28-91). The median number of patients without the target condition--needed to determine a test's specificity--was 76 (27-209). Two of the 43 studies (5%) reported a priori calculations of sample size. Twenty articles (47%) reported results for patient subgroups. The number of subgroups ranged from two to 19 (median four). No studies reported that sample size was calculated on the basis of preplanned analyses of subgroups. CONCLUSION: Few studies on diagnostic accuracy report considerations of sample size. The number of participants in most studies on diagnostic accuracy is probably too small to analyse variability of measures of accuracy across patient subgroups.

Confidence Intervals↗

Sampling Asian minorities to assess health and welfare.

STUDY OBJECTIVE: The aims were (1) to sample a specified subgroup of the Asian minority; (2) to give proper representation to those outside the areas of concentration; and (3) to evaluate the costs and benefits of the method. DESIGN: Glasgow postcodes with varying concentrations of Asians were sampled, and 173 Asians aged 30-40 were interviewed after household screening of 1439 Asian names identified on the electoral roll or valuation roll. Areas with few Asians, and households with two or more members aged 30-40, were undersampled, and then reweighted. MEASUREMENTS AND MAIN RESULTS: Nurse measures of blood pressure, lung function, and body mass were taken, and selected interview measures of health and social background are reported. Substantial differences in blood pressure, reported health, and social background were revealed between Asians in areas of concentration and those in areas of dispersion. Loss in effective sample size due to undersampling and reweighting was 4-5% in the case of the area sampling, 13% in the case of the household sampling. Losses of potential sample members through under registration were probably less than 6%. CONCLUSIONS: The present sampling method targets subgroups successfully, and improves on sampling in areas of concentration, in that it enables dispersed members of the minority, who differ in crucial indices of health and social position, to be represented. The costs of the method are acceptable.

Adolescent↗