Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Binomial Distribution”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 829 records · Page 46Linked to original sources

Simple stochastic models and their power-law type behaviour.

A power-law relationship between the mean and variance of ecological time series has been shown to hold for a vast number of species. Here we examine the behaviour of single-species stochastic models and concentrate in particular on the mean-variance relationship as the carrying capacity becomes large. Single-species stochastic models can be written as Markov chains, and the long-term distribution of population sizes and hence power-law scaling can be found analytically. The various power-law scalings that arise have very different biological implications for the effects of stochasticity and the departure from the deterministic paradigm. Finally we extend our analysis to consider the complicating factors of spatial heterogeneity, nontrivial deterministic dynamics, and multispecies models.

Bias↗

Statistical inference in generalized linear mixed models: a review.

We present a review of statistical inference in generalized linear mixed models (GLMMs). GLMMs are an extension of generalized linear models and are suitable for the analysis of non-normal data with a clustered structure. A GLMM contains parameters common to all clusters (fixed regression effects and variance components) and cluster-specific parameters. The latter parameters are assumed to be randomly drawn from a population distribution. The parameters of this population distribution (the variance components) have to be estimated together with the fixed effects. We focus on the case in which the cluster-specific parameters are normally distributed. The cluster-specific effects are integrated out of the likelihood so that the fixed effects and variance components can be estimated. Unfortunately, the integral over the cluster-specific effects is intractable for most GLMMs with a normal mixing distribution. Within a classical statistical framework, we distinguish between two broad classes of methods to handle this intractable integral: methods that rely on a numerical approximation to the integral and methods that use an analytical approximation to the integrand. Finally, we present an overview of available methods for testing hypotheses about the parameters of GLMMs.

Analysis of Variance↗

Computational simulations of vocal fold vibration: Bernoulli versus Navier-Stokes.

The use of the mechanical energy (ME) equation for fluid flow, an extension of the Bernoulli equation, to predict the aerodynamic loading on a two-dimensional finite element vocal fold model is examined. Three steady, one-dimensional ME flow models, incorporating different methods of flow separation point prediction, were compared. For two models, determination of the flow separation point was based on fixed ratios of the glottal area at separation to the minimum glottal area; for the third model, the separation point determination was based on fluid mechanics boundary layer theory. Results of flow rate, separation point, and intraglottal pressure distribution were compared with those of an unsteady, two-dimensional, finite element Navier-Stokes model. Cases were considered with a rigid glottal profile as well as with a vibrating vocal fold. For small glottal widths, the three ME flow models yielded good predictions of flow rate and intraglottal pressure distribution, but poor predictions of separation location. For larger orifice widths, the ME models were poor predictors of flow rate and intraglottal pressure, but they satisfactorily predicted separation location. For the vibrating vocal fold case, all models resulted in similar predictions of mean intraglottal pressure, maximum orifice area, and vibration frequency, but vastly different predictions of separation location and maximum flow rate.

Binomial Distribution↗

Multipoint mapping under genetic interference.

Genetic chiasma interference occurs when one crossover influences the probability of another crossover occurring nearby. While interference is known to occur in humans, it is typically ignored when computing multipoint likelihoods for genetic mapping. This biologically unsound assumption of no interference facilitates the calculation of the likelihoods at the expense of reduced power to accurately construct a genetic map. We have developed a computer program that calculates multipoint likelihoods of three-generation nuclear families while taking interference into account. In our program, interference is modelled by using a map function to convert genetic distances into recombination fractions. We can determine which of several map functions best fits the data by comparing the multipoint likelihoods of the data under each map function. Since the distribution of the difference between likelihoods is unknown, we use a simulation approach to determine the statistical significance of our results. When our program is applied to six loci, D10S34, D10S19, D10S16, D10S14, D10S4, and D10S20, from the CEPH consortium map of chromosome 10, we find significant evidence in favor of positive interference as modelled by the Sturt map function.

Binomial Distribution↗

Resampling techniques in the analysis of non-binormal ROC data.

The methods most commonly used for analyzing receiver operating characteristic (ROC) data incorporate "binormal" assumptions about the latent frequency distributions of test results. Although these assumptions have proved robust to a wide variety of actual frequency distributions, some data sets do not "fit" the binormal model. In such cases, resampling techniques such as the jackknife and the bootstrap provide versatile, distribution-independent, and more appropriate methods for hypothesis testing. This article describes the application of resampling techniques to ROC data for which the binormal assumptions are not appropriate, and suggests that the bootstrap may be especially helpful in determining confidence intervals from small data samples. The widespread availability of ever-faster computers has made resampling methods increasingly accessible and convenient tools for data analysis.

Bias↗

Multivariate Bayesian analysis of Gaussian, right censored Gaussian, ordered categorical and binary traits using Gibbs sampling.

A fully Bayesian analysis using Gibbs sampling and data augmentation in a multivariate model of Gaussian, right censored, and grouped Gaussian traits is described. The grouped Gaussian traits are either ordered categorical traits (with more than two categories) or binary traits, where the grouping is determined via thresholds on the underlying Gaussian scale, the liability scale. Allowances are made for unequal models, unknown covariance matrices and missing data. Having outlined the theory, strategies for implementation are reviewed. These include joint sampling of location parameters; efficient sampling from the fully conditional posterior distribution of augmented data, a multivariate truncated normal distribution; and sampling from the conditional inverse Wishart distribution, the fully conditional posterior distribution of the residual covariance matrix. Finally, a simulated dataset was analysed to illustrate the methodology. This paper concentrates on a model where residuals associated with liabilities of the binary traits are assumed to be independent. A Bayesian analysis using Gibbs sampling is outlined for the model where this assumption is relaxed.

Bayes Theorem↗

On selecting markers for association studies: patterns of linkage disequilibrium between two and three diallelic loci.

Association studies depend on linkage disequilibrium (LD) between a causative mutation and linked marker loci. Selecting markers that give the best chance of showing useful levels of LD with the causative mutation will increase the chances of successfully detecting an association. This report examines the variation in the extent of LD between a disease locus and one or two diallelic marker loci (termed single nucleotide polymorphisms or SNPs). We use a simulation method based on the neutral coalescent in a population of variable size to find the distribution of LD as a function of allele frequencies, the recombination rate, and the population history. Given that LD exists, the allele frequencies determine if a site will be useful for detecting an association with the disease mutation. We show that there is extensive variation in LD even for closely linked loci, implying that several markers may be needed to detect a disease locus. The distribution of LD between common variants is strongly influenced by ancestral population size. We show that in general, best results will be obtained if the frequencies of marker alleles are at least as large as the frequency of the causative mutation. Haplotypes of two or more SNPs generally have a higher probability than individual SNPs of showing useful LD with a disease mutation, although exceptions are described.

Alleles↗

Modification of an ELISA-based procedure for affinity determination: correction necessary for use with bivalent antibody.

A recently described procedure for the evaluation of the affinity of monoclonal antibodies [Friguet et al., J. Immun. Meth. 77, 305-319 (1985)] uses an ELISA system to determine the quantity of free antibody present in a mixture of antigen and antibody. However, an intact IgG may bind antigen by either of two binding sites, and an IgG can bind to a solid-phase antigen whether one or two of its binding sites are free. Therefore, this procedure does not directly provide the concn of liganded binding sites, the quantity necessary for calculation of the thermodynamic association constant. A binomial probability distribution relates the fraction of liganded binding sites to the concn of unliganded, singly liganded, and doubly liganded IgG assuming that the binding of each Fab to antigen is independent. Simulated experiments were used to compare the apparent binding characteristics of bivalent IgG and monovalent Fab and to calculate apparent association constants in each case. It was found that the affinity of binding sites on intact IgG was underestimated by a factor of at least 2 and that the error was inversely related to the fraction of liganded binding sites. Binding site affinity of an antibody may be underestimated by several orders of magnitude. On the basis of binomial analysis, it is possible to convert apparent concns of bound IgG to actual concns of liganded binding site resulting in the calculation of valid association constants for intact IgG without alteration of the experimental protocol.

Antibody Affinity↗

Exact group-sequential designs for clinical trials with randomized play-the-winner allocation.

The use of both sequential designs and adaptive treatment allocation are effective in reducing the number of patients receiving an inferior treatment in a clinical trial. In large samples, when the asymptotic normality of test statistics can be utilized, a standard sequential design can be combined with adaptive allocation. In small samples the planned error rate constraints may not be satisfied if normality is assumed. We address this problem by constructing sequential stopping rules with specified properties by consideration of the exact distribution of test statistics under a particular adaptive allocation scheme, the randomized play-the-winner rule. Using this approach, compared to traditional equal allocation trials, trials with adaptive allocation are shown to require a larger total sample size to achieve a given power. More interestingly, the expected number patients allocated to the inferior treatment may also be larger for the adaptive allocation designs depending on the true success rates.

Binomial Distribution↗

Strong intrinsic biases towards mutation and conservation of bases in human IgVH genes during somatic hypermutation prevent statistical analysis of antigen selection.

Immunoglobulin V region genes acquire point mutations during affinity maturation of the T-cell-dependent B-cell response. It has been proposed that both selection by antigen and characteristics of the DNA sequence are involved in determining the distribution of mutations along the genes. There is a tendency for replacement mutations to occur in the complementarity-determining regions and for silent mutations to accumulate in the framework regions of used genes. By analysing a group of highly mutated human IgVH4-34 (VH4. 21) and family 5 genes derived from human gut-associated lymphoid tissues, which were out-of-frame between VH and JH (and therefore not used) we have investigated the distribution of mutations acquired in the absence of selection. We observed that these genes may show the statistical hallmarks of selected genes, suggesting that intrinsic biases alone may be enough to give the appearance of selection. These data suggest that analysis of the distribution of mutations in IgVH genes cannot be used reliably to state whether antigenic selection of the B-cell carrying the genes occurred. In-frame genes had more silent mutations than the out-of-frame genes and lacked stop codons. These characteristics were considered to be indicative of selection in the in-frame genes derived from human gut-associated lymphoid tissue.

Antibody Affinity↗

Populations of acanthocephalus anguillaePomphorhynchus laevis in rivers with different pollution levels

The distribution of two acanthocephalan species (Pomphorhynchus laevisAcanthocephalus anguillae) in the chub (Leuciscus cephalus) was studied in four river reaches characterized by different levels of pollution: the River Ticino near Abbiategrasso (unpolluted), the Naviglio Grande Canal, in Milano (slightly polluted), the River Lambro near Merone village (polluted) and the River Lambro near Monza (severely polluted).Pomphorhynchus laevis was restricted to the unpolluted and the slightly polluted sites, while the intensity of A. anguillae increased proportionally to water pollution. These differences were partially explained by the variation in abundance of their intermediate hosts (Echinogammarus stammeri for P. laevisAsellus aquaticus for A. anguillae). Data on the occurrence of P. laevis and A. anguillae showed a significant negative binomial frequency distribution, suggesting their tendency to be aggregated within the host populations of L. cephalus.

Journal Article↗

Ongoing hypermutation in the Ig V(D)J gene segments and c-myc proto-oncogene of an AIDS lymphoma segregates with neoplastic B cells at different sites: implications for clonal evolution.

To investigate the role of somatic Ig hypermutation in the evolution of AIDS-associated B cell lymphomas, we analyzed the Ig V(D)J and c-myc genes expressed by neoplastic B cells in two extranodal sites, testis and orbit, and clonally related cells in the bone marrow. Testis and orbit B cells expressed differentially mutated but collinear V(H)DJ(H), V kappa J kappa and c-myc gene sequences. Shared mutations accounted for 10.2%, 8.4%, and 4.3% of the overall V(H)DJ(H), V kappa J kappa, and c-myc gene sequences. Tumor-site specific V(H)DJ(H), V kappa J kappa, and c-myc mutations were comparable in frequency, and a single point-mutation gave rise to an EcoRI site in the testis c-myc DNA. Both shared and tumor site-specific V(H)DJ(H), V kappa J kappa, and c-myc mutations displayed predominance of transitions over transversions. The "neoplastic" V(H)DJ(H) sequence was expressed by about 10(-5) cells in the bone marrow, and contained two of the three orbital, but none of the testicular V(H)DJ(H) mutations. The nature and distribution of the Ig V(D)J mutations found in the kappa chain suggested a selection by antigen in testis and orbit. Our data suggest that, in AIDS-associated B cell lymphomas, the Ig hypermutation machinery targets V(H)DJ(H), V kappa J kappa, and c-myc genes with comparable efficiency and modalities.

Adult↗

A Bayesian approach to logistic regression models having measurement error following a mixture distribution.

To estimate the parameters in a logistic regression model when the predictors are subject to random or systematic measurement error, we take a Bayesian approach and average the true logistic probability over the conditional posterior distribution of the true value of the predictor given its observed value. We allow this posterior distribution to consist of a mixture when the measurement error distribution changes form with observed exposure. We apply the method to study the risk of alcohol consumption on breast cancer using the Nurses Health Study data. We estimate measurement error from a small subsample where we compare true with reported consumption. Some of the self-reported non-drinkers truly do not drink. The resulting risk estimates differ sharply from those computed by standard logistic regression that ignores measurement error.

Age Factors↗

Distribution and sampling of the whiteflies Aleurothrixus floccosus, Dialeurodes citri, and Parabemisia myricae (Homoptera: Aleyrodidae) in citrus in Spain.

From 1993 to 1995 data sets were collected from four citrus groves in Valencia, Spain, to determine the distribution patterns of eggs and nymphs of Aleurothrixus floccosus (Maskell), Dialeurodes citri (Ashmead), and Parabemisia myricae (Kuwana) on leaves, and to develop reliable sampling plans for estimating densities of immature whiteflies. A. floccosus showed higher aggregation than the other two species. The dispersion index b from the Taylor power law did not vary between different developing stages for A. floccosus and D. citri, reaching overall values of 1.70 and 1.53, respectively. In P. myricae, b was 1.60 for eggs and N1, and 1.46 for the remaining nymphs. The minimum number of leaves to estimate the population density with a coefficient of variation of 0.25 for densities above 10 immature whiteflies per leaf was 40 for D. citri and P. myricae, and 250 for A. floccosus. Binomial sampling programs for the three species were rejected for pest management purposes due to the high sample sizes required. The enumerative procedure of counting the number of insects per leaf appears to be the most suitable method for D. citri and P. myricae. For A. floccosus an index of occupation (from 0 to 10) linearly related to the proportion of the leaf undersurface occupied by this insect was found to be reliable and time-saving. Examining 150 leaves with this index achieves the desired relative variation level of 0.25 for most population densities usually found in commercial groves.

Animals↗

The role of HLA genes in familial spondyloarthropathy: a comprehensive study of 70 multiplex families.

OBJECTIVES: To investigate whether HLA alleles, other than HLA-B27, influence predisposition to spondyloarthropathy (SpA) in multiplex families. METHODS: Seventy French families with at least two affected SpA members were recruited. Patients, and their first degree relatives were typed for HLA-A, B, C, and DR, and extended HLA haplotypes were determined. The distribution of HLA-A, C, and DR alleles carried on HLA-B27+ haplotypes in SpA families was compared with the distribution of these alleles among HLA-B27+ haplotypes in the French general population. Contribution to SpA susceptibility of HLA-A, B, C, and DR alleles, other than HLA-B27, was tested by transmission disequilibrium test. The contribution of HLA alleles to specific presentation features of SpA was examined. RESULTS: Frequencies of HLA-A, C, and DR alleles carried on HLA-B27+ haplotypes from SpA families were comparable with those seen in the French population, except for DR13 which was overrepresented among patients (pcorr<0.001). Most interestingly, the HLA-DR4 allele was transmitted in excess to patients with SpA, independently of linkage to HLA-B27 (pcorr=0.05), and in a direction opposite to that for HLA-B27+ unaffected siblings (pcorr=0.01). Finally, the distribution of HLA alleles was not related to the presentation feature of SpA. CONCLUSION: HLA predisposition to familial SpA appears not to be limited to HLA-B27, but some HLA-DR alleles also have a significant influence. In particular, HLA-DR4 contributes significantly to a genetic predisposition to SpA, which may have implications in our understanding of SpA pathogenesis.

Adult↗

Modelling the prevalence of Echinococcus and Taenia species in small ruminants of different ages in northern Jordan.

A base-line survey of Echinococcus granulosus, Taenia hydatigena and T ovis were undertaken in order to investigate the transmission dynamics of these parasites in northern Jordan. Intensity of E. granulosus infection, in sheep, increased with age in a linear fashion whilst the asymptotic prevalence was one. This implied that E. granulosus is in an endemic steady state with no evidence of protective immunity in the intermediate host. The mean number of cysts increased by 1.66 per year with approximately 0.320 infections per year, each infection consisting of 598 eggs to produce 5.2 cysts. The basic reproduction ratio (R0) was estimated to be 1.5-1.8. A similar pattern was suggested with E. granulosus in goats but the infection pressure appeared to be lower with only 0.128 cysts per year. Although infection in goats appeared to be endemic there was some evidence of departure from the model which might indicate that the model needs adjusting for this species. In the case of T. hydatigena the host age-intensity helminth distribution indicated that this parasite was hyperendemic in both sheep and goats, implying regulation by intermediate host immunity. Consequently, R0 was determined from asymptotic prevalence curves for T hydatigena and was calculated to be 4.0 and 3.1 in sheep and goats, respectively. The lower R0 in goats, together with the higher asymptotic age-intensity and age-prevalences, indicates that goats acquire immunity more slowly to T hydatigena in comparison to sheep. Taenia ovis was not detected in any animals.

Animals↗

Forecasting herd structure and milk production for production risk management.

Substantial increases in milk price volatility have resulted from changes in federal dairy policies. For a dairy farm, however, monthly gross milk receipts are a function of unit price and quantity produced. Both can vary substantially over time. Therefore, to be effective, risk management strategies must address milk and input price volatility (price risk management) and fluctuations in milk production per cow and cow numbers (production risk management). Herd milk production through time can be modeled as a discrete stochastic process using finite Markov chains. Cows at time t = 0 are assigned to homogeneous production cells in four-dimensional arrays with coordinates determined by parity (1, 2, 3), week in milk (1, ..., 104), pregnancy status (0, 1), and week pregnant (1, ..., 40). The processes of aging, pregnancy, involuntary cull, voluntary cull, abortion, dry-off, and freshening from week i-1 to week i are accounted for, using nonstationary transition probabilities. Bayesian estimates of transition probabilities are derived from historical herd data, assuming that individual outcomes are from Bernoulli distributions. The values of parameters theta(i) for the Bernoulli distributions are unknown but have prior distributions that follow beta distributions with parameters alpha(i) and beta(i) estimated from historical data. Herd observations are then used to generate posterior distributions of theta(i), also from beta distributions. Projecting from one week to the next is accomplished by moving virtual animals from one production cell to the next based on the transition probability assigned to that path. Summing production estimates and variances of all independent cells provides for an expected herd production with an associated variance. As expected, the forecast variance increases with time, reflecting increased uncertainty of distant projections. Model validation presents an interesting problem because future observations used for validation are under human control and are not independent of the forecast.

Animal Husbandry↗

A review of some extensions to generalized linear models.

Although generalized linear models are reasonably well known, they are not as widely used in medical statistics as might be appropriate, with the exception of logistic, log-linear, and some survival models. At the same time, the generalized linear modelling methodology is decidedly outdated in that more powerful methods, involving wider classes of distributions, non-linear regression, censoring and dependence among responses, are required. Limitations of the generalized linear modelling approach include the need for the iterated weighted least squares (IWLS) procedure for estimation and deviances for inferences; these restrict the class of models that can be used and do not allow direct comparisons among models from different distributions. Powerful non-linear optimization routines are now available and comparisons can more fruitfully be made using the complete likelihood function. The link function is an artefact, necessary for IWLS to function with linear models, but that disappears once the class is extended to truly non-linear models. Restricting comparisons of responses under different treatments to differences in means can be extremely misleading if the shape of the distribution is changing. This may involve changes in dispersion, or of other shape-related parameters such as the skewness in a stable distribution, with the treatments or covariates. Any exact likelihood function, defined as the probability of the observed data, takes into account the fact that all observable data are interval censored, thus directly encompassing the various types of censoring possible with duration-type data. In most situations this can now be as easily used as the traditional approximate likelihood based on densities. Finally, methods are required for incorporating dependencies among responses in models including conditioning on previous history and on random effects. One important procedure for constructing such likelihoods is based on Kalman filtering.

Animals↗