Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Binomial Distribution”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 865 records · Page 48Linked to original sources

Estimation of time since infection using longitudinal disease-marker data.

We propose a method to estimate the usually unknown time since infection for individuals infected with human immunodeficiency virus type 1 (HIV-1). If we assume the time since infection has an exponential prior distribution, then under the model the conditional distribution of time since infection, given the CD4 level at the time of the first positive HIV-1 antibody test, is a truncated normal density. We applied the method to prevalent cohort data both from intravenous drug users and from homosexual/bisexual men. For the intravenous drug users the estimated mean time since infection was 15.0 months from infection at a presumed mean CD4 level of 1060 cells/ml to first positive antibody test at a CD4 level of 597 cells/ml, which was the average CD4 at enrollment for infected subjects. For the homosexual/bisexual men the estimated mean time since infection was 16.7 months from infection at a presumed mean CD4 level of 699 cells/ml to first positive antibody test at an average CD4 level of 577 cells/ml. We performed a validation study using initially seronegative subjects in these cohorts who seroconverted to HIV-1-positive antibody status during the follow-up period. For the intravenous drug users, data were too few to provide definitive verification of the method. In the cohort of homosexual/bisexual men, however, there was a total of 70 seroconverters with relevant data. Among them, the median absolute difference between the midpoint of the known seroconversion interval and the estimated mean infection date was 4.6 months, conditional on CD4-lymphocyte measurements taken approximately 18 months subsequent to infection. Conditional on CD4 approximately 30 months after infection, this median difference increased modestly to 8.2 months. Our analysis suggested that the underlying mathematical model tends to overestimate short times since infection and underestimate long times since infection. We consider potential corrective modifications to the model.

Bias↗

Chi 2 tests: how useful are they in the analysis of medical research data?

This paper has outlined analyses of data for which the chi 2 test is commonly applied and has shown that on its own, a chi 2 test provides very little information about the interpretation of the data. In goodness to fit tests, at best, the chi 2 test is a mere first step in the interpretation of the data and much more information can be gleaned from the deviations between the observed and theoretical distributions, and graphical methods are more sensitive for assessing Normality. Similarly, the chi 2 test as the analysis of a 2 x 2 contingency table misses the most important features of the data. A chi 2 value and associated p value should not be presented without an estimation of the effect and its confidence interval. The use of NS, (not significant), after a chi 2 value (or any other test statistic) is particularly meaningless as it does not even specify a level of significance. The analysis of 2 x 2 contingency tables should specify a measure of the effect such as the difference between two proportions, their ratio or the odds ratio. Confidence intervals should be calculated and least importantly, statistical significance assessed by the Standard Normal Deviate. Chi 2 has no useful role to play in the analysis of 2 x 2 contingency tables. In fact, even with larger two dimensional tables, although chi 2 can be used to test the significance of associations between the rows and columns, such results are seldom of much interest as they give no indication of where in the table associations may exist. They are insensitive particularly if the categories of the row or column variables are quantitative since chi 2 takes no account of the ordering of the rows and columns. Although simple analyses of such tables are inefficient the advent of desk-top computers means that more sophisticated techniques such as logistic regression for proportions and Poisson regression for larger or multidimensional contingency tables can be applied. Although these methods may involve chi 2 tests for testing the significance of including or excluding variables, they are always associated with estimates of effects as measured by regression coefficients. The use of chi 2 for the simple analysis of contingency tables has no place in the interpretation of medical research data.

Binomial Distribution↗

P values.

Explore the source record for details and available documents.

Binomial Distribution↗

Formulae and tables for the determination of sample sizes and power in clinical trials for testing differences in proportions for the two-sample design: a review.

This paper is a compendium of exact and asymptotic formulae and tables for estimating the sample size in a clinical trial with two treatment groups and a dichotomous outcome. The paper provides separate formulae for equal and unequal treatment group sizes, formulae for the calculation of power given the sample size, and complete references for all formulae and tables cited.

Binomial Distribution↗

Assessment of small health risks based on exact sample sizes.

Exact sample sizes and critical numbers of cases for the rejection of a known event probability (10(-2) to 10(-6)) in favour of an increased probability (1.5- to 50-fold) at levels -alpha;beta- = -0.05; 0.10- and -alpha;beta- = -0.10;0.05- are presented. The numbers are thoroughly validated using the characteristics of the confidence interval for the unknown true event probability. Equivalence is shown to be obtainable for the tolerated maximal value of the relative risk and the upper limit of the confidence interval for the true event probability. Also demonstrated is the use of the tables for planned actions to reduce given empirical risks. In addition, use of the tables is shown for judging results from given data sets.

Bias↗

Limitations to the robustness of binormal ROC curves: effects of model misspecification and location of decision thresholds on bias, precision, size and power.

This paper concerns robustness of the binormal assumption for inferences that pertain to the area under an ROC curve. I applied the binormal model to rating method data sets sampled from bilogistic curves and observed small biases in area estimates. Bias increased as the range of decision thresholds decreased. The variance of area estimates also increased as the range of decision thresholds decreased. Together, minor bias and inflated variance substantially altered the size and power of statistical tests that compared areas under bilogistic ROC curves. I repeated the simulations by applying the binormal assumption to data sampled from binormal curves. Biases in area estimates were minimal for the binormal data, but the variance of area estimates was again higher when the range of decision thresholds was narrow. The size of tests that compared areas did not vary from the chosen significance level. Power fell, however, when the variance of area estimates was inflated. I conclude that inferences derived from the binormal assumption are sensitive to model misspecification and to the location of decision thresholds. A narrow span of decision thresholds increases the variability of area estimates and reduces the power of area comparisons. Model misspecification produces bias that alters test size and can exaggerate the loss of power that accompanies increased variability.

Bias↗

A non-iterative accurate asymptotic confidence interval for the difference between two proportions.

I propose a new confidence interval for the difference between two binomial probabilities that requires only the solution of a quadratic equation. The procedure is based one estimating the variance of the observed difference at the boundaries of the confidence interval, and uses least squares estimation rather than maximum likelihood as previously suggested. The proposed procedure is non-iterative, agrees with the conventional test of equality of two binomial probabilities, and, even for fairly small sample sizes, appears to yields actual 95 per cent confidence intervals with mean or median probabilities of coverage very close to 0.95. The Yates continuity correction appears to generate confidence intervals with the conditional probability of coverage at least equal to nominal levels.

Binomial Distribution↗

Interval estimation for the difference between independent proportions: comparison of eleven methods.

Several existing unconditional methods for setting confidence intervals for the difference between binomial proportions are evaluated. Computationally simpler methods are prone to a variety of aberrations and poor coverage properties. The closely interrelated methods of Mee and Miettinen and Nurminen perform well but require a computer program. Two new approaches which also avoid aberrations are developed and evaluated. A tail area profile likelihood based method produces the best coverage properties, but is difficult to calculate for large denominators. A method combining Wilson score intervals for the two proportions to be compared also performs well, and is readily implemented irrespective of sample size.

Binomial Distribution↗

Evaluating effects of exposures on embryo viability and uterine receptivity in in vitro fertilization.

We consider models for the occurrence of pregnancy following in vitro fertilization. In this clinical protocol, implantation depends on two factors: the receptivity of the uterus and the viability of at least one of the embryos transferred to the uterus. This work is motivated by the need to identify reliable bio-markers for these two factors, in order to enhance the success rate for couples undergoing this procedure. We present a general latent variable structure model for outcomes that take either one of two possible forms: as summed Bernoullis, based on an ultrasound count of gestational sacs, or as aggregated Bernoullis, based only on the outcome of a biochemical pregnancy test. We allow both uterine receptivity and embryo viability to be influenced by covariates. The proposed latent variable structure allows us to utilize the existing statistical packages to maximize an otherwise intractable likelihood function. The method is sufficiently flexible to permit any valid choice of link function. We illustrate by applying the method to a recent study of in vitro fertilization carried out in North Carolina. The number of cells at transfer is evidently a marker for embryo viability.

Adult↗

Adverse effects of chorionic villus sampling: a meta-analysis.

Meta-analysis is a popular tool for combining evidence from several related studies. The technique is usually used to combine randomized clinical trials, case-control studies or prospective studies where each study has its own exposed and unexposed groups. By including separate 'study effects' (either fixed or random), one can combine information about differences between control and exposed groups, while still allowing for study heterogeneity. In this paper, we extend existing methods to combine studies of disparate designs, where some studies do not include concurrent controls. We apply the methods to a meta-analysis of the association of prenatal testing via chorionic villus sampling with the occurrence of terminal transverse limb defects.

Binomial Distribution↗

Directly modelling matched case-control data.

Matching in case-control studies is a situation in which one wishes to make inferences about a parameter of interest in the presence of nuisance parameters. The usual approach is to apply a conditional likelihood. A bivariate latent class log-linear model for binomial responses is shown to yield a standard likelihood identical to the usual conditional one. This extension of the Rasch model for binary responses gives consistent estimates and a suitable likelihood function for cases matched with any fixed number of controls.

Age Factors↗

GEE analysis of negatively correlated binary responses: a caution.

The method of generalized estimating equations has become almost standard for analysing longitudinal and other correlated response data. However, we have found that if binary responses have less than binomial variation over clusters, and are modelled using exchangeable correlations, prevailing software implementations may give unreliable results. Bounding the negative correlation away from its theoretical minimum may not always be a satisfactory solution. In such instances, using the independence working correlation structure and robust SEs is a more trustworthy alternative.

Binomial Distribution↗

Genetic variance components analysis for binary phenotypes using generalized linear mixed models (GLMMs) and Gibbs sampling.

The common complex diseases such as asthma are an important focus of genetic research, and studies based on large numbers of simple pedigrees ascertained from population-based sampling frames are becoming commonplace. Many of the genetic and environmental factors causing these diseases are unknown and there is often a strong residual covariance between relatives even after all known determinants are taken into account. This must be modelled correctly whether scientific interest is focused on fixed effects, as in an association analysis, or on the covariances themselves. Analysis is straightforward for multivariate Normal phenotypes, but difficulties arise with other types of trait. Generalized linear mixed models (GLMMs) offer a potentially unifying approach to analysis for many classes of phenotype including multivariate Normal traits, binary traits, and censored survival times. Markov Chain Monte Carlo methods, including Gibbs sampling, provide a convenient framework within which such models may be fitted. In this paper, Bayesian inference Using Gibbs Sampling (a generic Gibbs sampler; BUGS) is used to fit GLMMs for multivariate Normal and binary phenotypes in nuclear families. BUGS is easy to use and readily available. We motivate a suitable model structure for Normal phenotypes and show how the model extends to binary traits. We discuss parameter interpretation and statistical inference and show how to circumvent a number of important theoretical and practical problems that we encountered. Using simulated data we show that model parameters seem consistent and appear unbiased in smaller data sets. We illustrate our methods using data from an ongoing cohort study.

Binomial Distribution↗

Economic depression and the use of physician services in Finland.

At the start of the 1990s, the economic situation in Finland deteriorated radically. During the depression (1991-93), health care expenditure decreased by about 10%, and was associated with considerable changes in Finnish health care. This paper reports studies of the determinants of use of physician services in Finland in the 1990s. The particular aim was to evaluate how utilization altered during the economic depression and during the changes in the health care system. Using econometric methods, an attempt was made to describe the changes in structure and level of utilization. The study was based on annual computer-assisted telephone interviews made during 1991-94. Visits to a doctor were analysed using a two-part model (logit and truncated negative binomial regression). Structural changes were tested by Chow-type tests and changes in the level of utilization by chronologically defined dummy variables for each year. The most significant changes (both in structure and level) occurred in the model explaining the number of visits (negative binomial regression) of chronically ill persons. Variables describing the continuity of care seem to be more important determinants of visits to a doctor than certain other availability and socioeconomic variables.

Adult↗

The probability of a cancer cluster due to chance alone.

We propose to use a very simple model to test whether a cancer cluster is due to chance alone. We focus on the acute childhood leukaemia cluster in Columbus, Ohio. In 1975, 12 leukaemia cases were observed in Columbus while the expected number is 6 cases per year. According to our simple model, the probability of such an occurrence, due to chance alone, is less than 1 per cent. However, if we divide the child population of the U.S.A. into 200 regions (each region having 200 000 children) then the probability that at least one region will see, in a given year, 12 or more cases is higher than 80 per cent. So in this sense the Columbus cluster could be attributed to chance alone. However, the probability that any of the 200 regions see 18 cases or more in a given year is almost 0. Thus, a cluster of 18 or more cases in a region of 200 000 children should be regarded as highly suspicious and should be investigated.

Acute Disease↗