Search PubMedSearch

SEARCH · Search PubMed

Results for “Statistical Bootstrap”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Phylogenomic subsampling and upsampling for efficient evolutionary analyses of big data.

Long runtimes, high memory demands, and reliance on high-performance computing impede phylogenomic analyses. We review a scalable phylogenomic subsampling with upsampling (PSU) framework to address this challenge, which reduces runtime and memory requirements by orders of magnitude. In PSU, small subsamples of sites from a concatenated alignment are analyzed, which are expanded by upsampling before inference, and the resulting inferences are aggregated to obtain evolutionary estimates. PSU harnesses the fact that the computational cost of maximum likelihood analysis is strongly influenced by the number of distinct site patterns in the concatenated alignment, whereas statistical power depends primarily on the amount of evolutionary information represented by the total number of sites and substitutions. By reducing the former while restoring the latter through upsampling, PSU can approximate many full-alignment analyses at substantially lower computational cost. Analysis of simulated and empirical datasets shows that PSU can accurately estimate bootstrap support values, select the optimal substitution model, test evolutionary hypotheses, and infer branch lengths, divergence times, and associated uncertainty measures. PSU also provides distributions of inferred clade support across independent subsamples, enabling detection of conflicting phylogenetic signals that may remain hidden in conventional bootstrap analysis of concatenated alignments. Automated tuning of subsample size, the number of subsamples, and the number of upsampling replicates make PSU practical. We suggest that PSU is a general approach for scalable phylogenomic inference using a broad range of statistical methods. By enabling analyses of genome-scale alignments on commodity hardware, PSU broadens research access and reduces environmental and infrastructural costs of big-data phylogenomics.

Phylogeny

Phylogenomic subsampling and upsampling for efficient evolutionary analyses of big data.

Long runtimes, high memory demands, and reliance on high-performance computing impede phylogenomic analyses. We review a scalable phylogenomic subsampling with upsampling (PSU) framework, in which small subsamples of sites from a concatenated alignment are expanded by upsampling before inference, and the resulting analyses are then aggregated to obtain evolutionary estimates. PSU harnesses the fact that the computational cost of maximum likelihood analysis is strongly influenced by the number of distinct site patterns in the concatenated alignment, whereas statistical power depends primarily on the amount of evolutionary information represented by the total number of sites and substitutions. By reducing the former while restoring the latter through upsampling, PSU can approximate many full-data analyses at substantially lower computational cost. Analysis of simulated and empirical datasets shows that PSU can accurately estimate bootstrap support values, select the optimal substitution model, test evolutionary hypotheses, and infer branch lengths, divergence times, and associated uncertainty measures, while reducing runtime and memory requirements by orders of magnitude. PSU also provides distributions of inferred clade support across independent subsamples, enabling detection of conflicting phylogenetic signals that may remain hidden in conventional bootstrap analysis. Automated tuning of subsample size, the number of subsamples, and the number of upsampling replicates make PSU practical across diverse datasets. We suggest that PSU is a general strategy for scalable phylogenomic inference using a broad range of statistical methods. By enabling analyses of genome-scale alignments on commodity hardware, PSU broadens research access and reduces environmental and infrastructural costs of big-data phylogenomics.

confidence limits

Model selection and the estimation of odds ratios in the presence of extraneous factors.

This paper deals with model selection and the estimation of odds ratios from cross-classified frequencies in the presence of extraneous factors. The odds ratio is estimated in different ways dependent on whether the extraneous factor is modelled as an effect modifier, a confounder, or neither. Routinely this choice is based on statistical tests of null hypotheses. By contrast, we propose selection of the model which is estimated to maximize the accuracy of the estimator of the odds ratio, on average. We demonstrate how a non-parametric bootstrap method can be used to carry out the selection, and illustrate the methodology using an example on use of oral contraceptives and myocardial infarction.

Adult

Inference of horizontal genetic transfer from molecular data: an approach using the bootstrap.

Inconsistencies in taxonomic relationships implicit in different sets of nucleic acid sequences potentially result from horizontal transfer of genetic material between genomes. A nonparametric method is proposed to determine whether such inconsistencies are statistically significant. A similarity coefficient is calculated from ranked pairwise identities and evaluated against a distribution of similarity coefficients generated from resampled data. Subsequent analyses of partial data sets, obtained by the elimination of individual taxa, identify particular taxa to which the significance may be attributed, and can sometimes help in distinguishing horizontal genetic transfer from inconsistencies due to convergent evolution or variation in evolutionary rate. The method was successfully applied to data sets that were not found to be significantly different with existing methods that use comparisons of phylogenetic trees. The new statistical framework is also applicable to the inference of horizontal transfer from restriction fragment length polymorphism distributions and protein sequences.

Animals

Estimation of error rates in discriminant analysis with selection of variables.

Accurate estimation of misclassification rates in discriminant analysis with selection of variables by, for example, a stepwise algorithm, is complicated by the large optimistic bias inherent in standard estimators such as those obtained by the resubstitution method. Application of a bootstrap adjustment can reduce the bias of the resubstitution method; however, the bootstrap technique requires the variable selection procedure to be repeated many times and is therefore difficult to compute. In this paper we propose a smoothed estimator that requires relatively little computation and which, on the basis of a Monte Carlo sampling study, is found to perform generally at least as well as the bootstrap method.

Algorithms

Stability and reproducibility of time structure in spontaneous behavior of male rats.

The computer pattern recognition system for the study of spontaneous rat behavior has allowed new analytical techniques which expand the definition of experimentally induced changes in behavior. As with any technique, the stability of the measures must be considered when evaluating overall sensitivity. This study evaluates the stability and reproducibility of three behavioral measures: a measure of the number of initiations of specific behavioral acts, a measure of the total time of each act, and a measure of behavioral time structure. Normal statistical parameters are used to evaluate the significance of changes detected using the first two measures, but the third measure utilizes K-functions, the bootstrap and ad hoc criteria to evaluate significance of observed changes. This study compares the stability of results from these three measures as applied to fourteen different groups of control Sprague-Dawley male rats. All three measures provided stable and reproducible results, but the measure of time structure, the K-function analysis, provided the greatest consistency. Behavior, particularly spontaneous behavior, has traditionally been perceived as being intrinsically variable. However, this study shows that the computer pattern recognition system and its analytical techniques provide stable and reproducible values that vary only a few percent.

Amphetamine

Assessing proportionality in the proportional odds model for ordinal logistic regression.

The proportional odds model for ordinal logistic regression provides a useful extension of the binary logistic model to situations where the response variable takes on values in a set of ordered categories. The model may be represented by a series of logistic regressions for dependent binary variables, with common regression parameters reflecting the proportional odds assumption. Key to the valid application of the model is the assessment of the proportionality assumption. An approach is described arising from comparisons of the separate (correlated) fits to the binary logistic models underlying the overall model. Based on asymptotic distributional results, formal goodness-of-fit measures are constructed to supplement informal comparisons of the different fits. A number of proposals, including application of bootstrap simulation, are discussed and illustrated with a data example.

Biometry

A new test for linkage in the presence of locus heterogeneity.

The detection of linkage in complex traits, although potentially of the greatest value, has proved very difficult. One reason may be the drastic effect that locus heterogeneity has on statistical power. We propose a new test for linkage in the presence of heterogeneity, based upon the sum of individual pedigree maximum lod scores, combined with a bootstrap method for estimating the null-hypothesis distribution. The technique is designed to exploit modern computer capability and to avoid reliance on asymptotic-distribution theory. Numerical comparisons indicate that for small pedigrees this new test can detect linkage with 30%-50% less data than are required by standard methods. A computer program for simulating the distribution and for performing the test of linkage is available from the authors.

Female

18S rDNA sequences and the holometabolous insects.

The Holometabola (insects with complete metamorphosis: beetles, wasps, flies, fleas, butterflies, lacewings, and others) is a monophyletic group that includes the majority of the world's animal species. Holometabolous orders are well defined by morphological characters, but relationships among orders are unclear. In a search for a region of DNA that will clarify the interordinal relationships we sequenced approximately 1080 nucleotides of the 5' end of the 18S ribosomal RNA gene from representatives of 14 families of insects in the orders Hymenoptera (sawflies and wasps), Neuroptera (lacewing and antlion), Siphonaptera (flea), and Mecoptera (scorpionfly). We aligned the sequences with the published sequences of insects from the orders Coleoptera (beetle) and Diptera (mosquito and Drosophila), and the outgroups aphid, shrimp, and spider. Unlike the other insects examined in this study, the neuropterans have A-T rich insertions or expansion regions: one in the antlion was approximately 260 bp long. The dipteran 18S rDNA evolved rapidly, with over 3 times as many substitutions among the aligned sequences, and 2-3 times more unalignable nucleotides than other Holometabola, in violation of an insect-wide molecular clock. When we excluded the long-branched taxa (Diptera, shrimp, and spider) from the analysis, the most parsimonious (minimum-length) trees placed the beetle basal to other holometabolous orders, and supported a morphologically monophyletic clade including the fleas+scorpionflies (96% bootstrap support). However, most interordinal relationships were not significantly supported when tested by maximum likelihood or bootstrapping and were sensitive to the taxa included in the analysis. The most parsimonious and maximum-likelihood trees both separated the Coleoptera and Neuroptera, but this separation was not statistically significant.(ABSTRACT TRUNCATED AT 250 WORDS)

Animals

Phylogeny of some ascaridoid nematodes, inferred from comparison of 18S and 28S rRNA sequences.

Reverse transcription of cellular RNA was used to obtain sequences from regions of 18S and 26/28S ribosomal RNA for eight species of ascaridoid nematodes. Phylogenetic relationships among these species were inferred from the aligned sequences by maximum-parsimony and maximum-likelihood methods. Seventy-nine of the 168 sites that varied were phylogenetically informative in parsimony analysis. Phylogenetic inference based on maximum-likelihood analysis of all sequence sites yielded a tree of topology similar to that of the parsimony result. Monophyletic groups that were strongly supported by bootstrap resampling of these data included species constituting the Ascaridinae, as well as those representing the Toxocarinae. Alternative topologies that included a member of the Ascaridinae with the Toxocarinae were rejected statistically on the basis of analysis of the mean and variance of parsimony step differences between trees. The conformance of these sequence data to a molecular-clock model of evolution was evaluated statistically by a maximum-likelihood approach. The inferred rate of rRNA sequence change along the branch leading to Parascaris equorum is not consistent with a clocklike model of evolution.

Animals

Bootstrapped confidence intervals for the Cox model using a linear relative risk form.

A linear relative risk form for the Cox model is sometimes more appropriate than the usual exponential form. The usual asymptotic confidence interval may not have the appropriate coverage, however, due to flatness of the likelihood in the neighbourhood of beta. For a single continuous covariate, we derive bootstrapped confidence intervals with use of two resampling methods. The first resamples the original data and yields both one-step and fully iterated estimates of beta. The second resamples the score and information quantities at each failure time to yield a one-step estimate. We computed the bootstrapped confidence intervals by three different methods and compared these intervals to one based on the asymptotic standard error and to a likelihood-based interval. The bootstrapped intervals did not perform well and underestimated the true coverage in most cases.

Computer Simulation

A set of viral DNA decamers enriched in transcription control signals.

We studied the frequency distribution of oligonucleotides 10 bp long in a sample of 620 Kb of viral genomes, containing 102 sequences from GenBank, with the aim of detecting transcription control signals. Two thousand three hundred decamers had a frequency 10 times higher than the mean and were subjected to further statistical analysis. For each of the 2300 decamers (parents), we counted the individual frequencies of the 30 decamers differing from the parent by one base mutation (progeny) and then calculated two variance/mean chi squares for the progeny, with and without the parent. We then studied the distribution of the ratio between the two chi squares. Out of 2300 decamers, 10 times more frequent than average, 479 decamers had a chi square ratio of 1.9 or larger. In this final set, which corresponds to less than 0.05% of all possible decamers, 58 decamers were found to contain viral and eukaryotic transcription control elements, like NF-kB, Sp1 and others. Furthermore, this set contains an excess of signals of length 5, 6, 7, 8, 9 and 10, when compared to 150 random sets, bootstrapped from the same viral genomes.

Algorithms

[Longitudinal study of the evolution of the frequency of dental caries in a school milieu: a statistical model].

Following a 3-year epidemiological study of the appearance of dental caries in a population of children aged 6 to 9 years (and examined every 12 months), a model of the evolution of the frequency of caries based on a Poisson "with zeros" distribution is proposed. The multivariate distribution of the above phenomena appears as a mixture of multiple negative binomial distribution (MNB) of the following form: sigma mjMNB(a;k0jb,k1jb,k2jb) with sigma mj = 1. The experimental data from the sample of 501 children validate the model in its successive stages. Simulations using the bootstrap method show that the estimates of the model parameters remain stable in the neighbourhood of the distribution observed; at the same time they permit establishing the margins of confidence. Finally, it is shown why the total number of dental faces affected during the experimental period remained strongly asymmetrical.

Child

Sampling strategies for toxic air contaminants.

Annual concentrations of toxic air contaminants are of primary concern from the perspective of chronic human exposure assessment and risk analysis. Despite recent advances in air quality monitoring technology, resource and technical constraints often impose limitations on the availability of a sufficient number of ambient concentration measurements for performing environmental risk analysis. Therefore, sample size limitations, representativeness of data, and uncertainties in the estimated annual mean concentration must be examined before performing quantitative risk analysis. In this paper, we discuss several factors that need to be considered in designing field-sampling programs for toxic air contaminants and in verifying compliance with environmental regulations. Specifically, we examine the behavior of SO2, TSP, and CO data as surrogates for toxic air contaminants and as examples of point source, area source, and line source-dominated pollutants, respectively, from the standpoint of sampling design. We demonstrate the use of bootstrap resampling method and normal theory in estimating the annual mean concentration and its 95% confidence bounds from limited sampling data, and illustrate the application of operating characteristic (OC) curves to determine optimum sample size and other sampling strategies. We also outline a statistical procedure, based on a one-sided t-test, that utilizes the sampled concentration data for evaluating whether a sampling site is compliance with relevant ambient guideline concentrations for toxic air contaminants.

Air Pollutants

Bootstrapping: a tool for clinical research.

The use of the bootstrap sampling technique is applied to the type of data found in clinical research. Confidence intervals are computed for simulated values by use of SAS. By applying this approach, clinical researchers are free to explore topics that do not meet the requirements of traditional statistical analytic methods.

Acquired Immunodeficiency Syndrome