Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Statistical Distributions”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,495 records · Page 83Linked to original sources

Efficient estimation of graphlet frequency distributions in protein-protein interaction networks.

MOTIVATION: Algorithmic and modeling advances in the area of protein-protein interaction (PPI) network analysis could contribute to the understanding of biological processes. Local structure of networks can be measured by the frequency distribution of graphlets, small connected non-isomorphic induced subgraphs. This measure of local structure has been used to show that high-confidence PPI networks have local structure of geometric random graphs. Finding graphlets exhaustively in a large network is computationally intensive. More complete PPI networks, as well as PPI networks of higher organisms, will thus require efficient heuristic approaches. RESULTS: We propose two efficient and scalable heuristics for finding graphlets in high-confidence PPI networks. We show that both PPI and their model geometric random networks, have defined boundaries that are sparser than the 'inner parts' of the networks. In addition, these networks exhibit 'uniformity' of local structure inside the networks. Our first heuristic exploits these two structural properties of PPI and geometric random networks to find good estimates of graphlet frequency distributions in these networks up to 690 times faster than the exhaustive searches. Our second heuristic is a variant of a more standard sampling technique and it produces accurate approximate results up to 377 times faster than the exhaustive searches. We indicate how the combination of these approaches may result in an even better heuristic. AVAILABILITY: Supplementary information is available at http://www.cs.toronto.edu/~natasha/BIOINF-2005-0946/Supplementary.pdf. Software implementing the algorithms is available at http://www.cs.toronto.edu/~natasha/BIOINF-2005-0946/estimate_grap-hlets.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Algorithms↗

Partitioning biochemical reference data into subgroups: comparison of existing methods.

Four existing methods for partitioning biochemical reference data into subgroups are compared. Two of these, the method of Sinton et al. and that of Ichihara and Kawai, are based on a quotient of a difference between the subgroups and the reference interval for the combined distribution. The criterion of Sinton et al. appears rather stringent and could lead to recommendations to apply a common reference interval in many cases where establishment of group-specific reference intervals would be more useful. The method of Ichihara and Kawai is similar to that of Sinton et al., but their criterion, based on a quantity derived from between-group and within-group variances, seems to lead to inconsistent results when applied to some model cases. These two methods have the common weakness of using gross differences between subgroup distributions as an indicator of differences between their reference limits, while distributions with different means can actually have equal reference limits and those with equal means can have different reference limits. The idea of Harris and Boyd to require that the proportions of the subgroup distributions outside the common reference limits be kept reasonably close to the ideal value of 2.5% as a prerequisite for using common reference limits seems to have been a major improvement. The other two methods considered, that of Harris and Boyd and the "new method" follow this idea. The partitioning criteria of Harris and Boyd have previously been shown to provide a poor correlation to those proportions, however, and the weaknesses of their method are summarized in a list of five drawbacks. Different versions of the new method offer improvements to these drawbacks.

Data Interpretation, Statistical↗

Unsupervised learning of a finite mixture model based on the Dirichlet distribution and its application.

This paper presents an unsupervised algorithm for learning a finite mixture model from multivariate data. This mixture model is based on the Dirichlet distribution, which offers high flexibility for modeling data. The proposed approach for estimating the parameters of a Dirichlet mixture is based on the maximum likelihood (ML) and Fisher scoring methods. Experimental results are presented for the following applications: estimation of artificial histograms, summarization of image databases for efficient retrieval, and human skin color modeling and its application to skin detection in multimedia databases.

Algorithms↗

Most recent common ancestor probability distributions in gene genealogies under selection.

A computational study is made of the conditional probability distribution for the allelic type of the most recent common ancestor in genealogies of samples of n genes drawn from a population under selection, given the initial sample configuration. Comparisons with the corresponding unconditional cases are presented. Such unconditional distributions differ from samples drawn from the unique stationary distribution of population allelic frequencies, known as Wright's formula, and are quantified. Biallelic haploid and diploid models are considered. A simplified structure for the ancestral selection graph of S. M. Krone and C. Neuhauser (1997, Theor. Popul. Biol. 51, 210-237) is enhanced further, reducing the effective branching rate in the graph. This improves efficiency of such a nonneutral analogue of the coalescent for use with computational likelihood-inference techniques.

Algorithms↗

Night-time quiescence and morning activation in the human colon: effect on transit of dispersed and large single unit formulations.

OBJECTIVES: Controlling the delivery of drugs to different regions of the colon remains an elusive goal. The aim of this study was to define the diurnal variation in colonic transit and show how this influences the colonic distribution and residence time of different formulations given either in the morning or evening. METHODS: Colonic transit of small particulates and a large capsule was measured during nocturnal sleep and daytime wakefulness. Eighteen healthy volunteers participated in a randomised crossover study. 111In-labelled resin (150-300 microm) and a large 99mTc-labelled non-disintegrating capsule (22 x 8 mm) were swallowed at either 0800h or 1700h. MAIN OUTCOME MEASURES: The geometric centre of isotope (range 1-9) was calculated from serial scintiscans allowing comparison of overnight and daytime transit. RESULTS: Transit of resin was delayed in the overnight compared to daytime 8 h periods (change in geometric centres (GCs), mean +/- SEM, 0.59 +/- 0.14 vs 1.46 +/- 0.39 respectively, P < 0.02). Maximal resin movement occurred immediately after awakening, prior to breakfast, in 9/18 studies (P < 0.05). The capsule was more distal than the resin at the end of the study 15 h after dosing (P < 0.001). There was marked inter-individual variability in distribution of both resin and capsule at 15 h, the range of GCs being 2.8-9 and 2.2-9, respectively. CONCLUSION: Sleep delays colonic transit and large capsules travel faster than dispersed small particles. However, substantial inter-individual variability in transit makes targeting specific regions of the human colon unreliable with either dispersed or single unit formulations.

Administration, Rectal↗

Relation between attempted suicide and suicide rates among young people in Europe.

STUDY OBJECTIVE: To determine if there are associations between rates of suicide and attempted suicide in 15-24 year olds in different countries in Europe. DESIGN: Attempted suicide rates were based on data collected in centres in Europe between 1989 and 1992 as part of the WHO/EURO Multicentre Study of Parasuicide. Comparison was made with both national suicide rates and local suicide rates for the areas in which the attempted suicide monitoring centres are based. SETTING: 15 centres in 13 European countries. PATIENTS: Young people aged 15-24 years who had taken overdoses or deliberately injured themselves and been identified in health care facilities. MAIN RESULTS: There were positive correlations (Spearman rank order) between rates of attempted suicide and suicide rates in both sexes. The correlations only reached statistical significance for male subjects: regional suicide rates, r = 0.65, p < 0.02; national suicide rates, r = 0.55, p < 0.02. CONCLUSIONS: Rates of attempted suicide and suicide in the young covary. The recent increase in attempted suicide rates in young male subjects in several European countries could herald a further increase in suicide rates.

Adolescent↗

Voltage noise influences action potential duration in cardiac myocytes.

Stochastic gating of ion channels introduces noise to membrane currents in cardiac muscle cells (myocytes). Since membrane currents drive membrane potential, noise thereby influences action potential duration (APD) in myocytes. To assess the influence of noise on APD, membrane potential is in this study formulated as a stochastic process known as a diffusion process, which describes both the current-voltage relationship and voltage noise. In this framework, the response of APD voltage noise and the dependence of response on the shape of the current-voltage relationship can be characterized analytically. We find that in response to an increase in noise level, action potential in a canine ventricular myocytes is typically prolonged and that distribution of APDs becomes more skewed towards long APDs, which may lead to an increased frequency of early after-depolarization formation. This is a novel mechanism by which voltage noise may influence APD. The results are in good agreement with those obtained from more biophysically-detailed mathematical models, and increased voltage noise (due to gating noise) may partially underlie an increased incidence of early after-depolarizations in heart failure.

Action Potentials↗

Stoichiometric design of metabolic networks: multifunctionality, clusters, optimization, weak and strong robustness.

Starting from a limited set of reactions describing changes in the carbon skeleton of biochemical compounds complete sets of metabolic networks are constructed. The networks are characterized by the number and types of participating reactions. Elementary networks are defined by the condition that a specific chemical conversion can be performed by a set of given reactions and that this ability will be lost by elimination of any of these reactions. Groups of networks are identified with respect to their ability to perform a certain number of metabolic conversions in an elementary way which are called the network's functions. The number of the network functions defines the degree of multifunctionality. Transitions between networks and mutations of networks are defined by exchanges of single reactions. Different mutations exist such as gain or loss of function mutations and neutral mutations. Based on these mutations neighbourhood relations between networks are established which are described in a graph theoretical way. Basic properties of these graphs are determined such as diameter, connectedness, distance distribution of pairs of vertices. A concept is developed to quantify the robustness of networks against changes in their stoichiometry where we distinguish between strong and weak robustness. Evolutionary algorithms are applied to study the development of network populations under constant and time dependent environmental conditions. It is shown that the populations evolve toward clusters of networks performing a common function and which are closely neighboured. Under changing environmental conditions multifunctional networks prove to be optimal and will be selected.

Algorithms↗

Bayesian fine-scale mapping of disease loci, by hidden Markov models.

We present a new multilocus method for the fine-scale mapping of genes contributing to human diseases. The method is designed for use with multiple biallelic markers-in particular, single-nucleotide polymorphisms for which high-density genetic maps will soon be available. We model disease-marker association in a candidate region via a hidden Markov process and allow for correlation between linked marker loci. Using Markov-chain-Monte Carlo simulation methods, we obtain posterior distributions of model parameter estimates including disease-gene location and the age of the disease-predisposing mutation. In addition, we allow for heterogeneity in recombination rates, across the candidate region, to account for recombination hot and cold spots. We also obtain, for the ancestral marker haplotype, a posterior distribution that is unique to our method and that, unlike maximum-likelihood estimation, can properly account for uncertainty. We apply the method to data for cystic fibrosis and Huntington disease, for which mutations in disease genes have already been identified. The new method performs well compared with existing multi-locus mapping methods.

Alleles↗

[Sonographically detectable changes in placental structures in pregnancy. 2. Statistical comparison of the frequency distribution of placenta stages 0-3 newborn infants with a birth weight of 2,500-3,999 grams].

1200 examinations of sonographical demonstrable placental ripeness were done in 552 pregnant women. The frequency of stages 0 to 3 as compared with aid of O. Bunke's confidence intervals. Stage 0 were found frequently until the 30. week of pregnancy, stage 1 between the 21. and 36. gestational week, stage 2 between the 33. and 38. week of pregnancy and stage 3 between the 35. and 40. week. The comparison of the areas of unsharpness several stages possibly may give informations about the period of pregnancy when the stages coincide.

Birth Weight↗

Is prenatal care really ineffective? Or, is the 'devil' in the distribution?

Prenatal care should improve infant health, yet research frequently finds only weak effects. If there are two kinds of pregnancies, 'complicated' and 'normal' ones, then combining these pregnancies may lead prenatal care to appear ineffective. Data from the National Maternal and Infant Health Survey (NMIHS) offers compelling evidence. The standard 2SLS approach yields obviously bimodal residuals and frequently insignificant prenatal care coefficients. In contrast, estimating birth weights with a finite mixture model yields estimates revealing that prenatal care has a substantial effect on 'normal' pregnancies. Our Monte Carlo experiment confirms that ignoring even a small proportion of 'complicated' pregnancies can lead prenatal care to appear unimportant.

Female↗

Creation of a low-risk reference group and reference interval of fasting venous plasma glucose.

Reference intervals are recommended for naturally occurring quantities and required in the evaluation of new components in order to provide clinically useful information. The aim of the present study is to present a method for selecting reference individuals for the determination of fasting venous plasma glucose (f-vPG) reference intervals and ways to determine if disease groups can share reference intervals with an ideal reference population. Reference subjects were randomly selected, eligibility was judged according to predetermined inclusion and exclusion criteria. Using the literature we selected risk indicators for diabetes mellitus (DM) and used these indicators to rule out high-risk individuals in order to obtain a reference distribution of f-vPG determined using individuals with low risk of DM. The distribution of f-vPG in the high-risk individuals was compared with that determined for the low-risk group. We then estimated the ability of the high-risk individuals to share the reference interval of the low-risk individuals, and calculated the fraction that was outside this interval. Distributions were also investigated for linearity in the cumulated frequency rankit distribution of In-values. The allowable difference between two reference limits could not exceed 0.375 times the population biological variation. Most risk indicators were powerful predictors of high f-vPG values. Subgroups with these risk indicators should not be included in the homogeneous In-normally distributed reference distribution. Distributions of f-vPG concentrations in individuals with risk factors were not homogeneous and varying percentages of individuals were outside the reference distribution, having f-vPG greater than 7.0 mmol/l. We conclude that randomisation is only useful to recruit candidate reference subjects. To rule out subjects according to clinical risk factors for diabetes, it is necessary to identify a reference population with low risk of exhibiting increased f-vPG concentrations. This method may be used to validate a reference interval for a particular analyte with respect to an investigated disease, and to stratify risk factors of importance.

Blood Glucose↗

The sampling distribution of kappa.

Research on Herrnstein's single-schedule equation contains conflicting findings; some laboratories report variations in the k parameter with reinforcer value, and others report constancy. The reported variation in k typically occurs across very low reinforcer values, and constancy applies across higher values. Here, simulations were conducted assuming a wide range of reinforcer values, and the parameters of Herrnstein's equation were estimated for simulated responding. In the simulations, responses controlled by current reinforcement contingencies were added to other responses ('noise'), controlled by the experimental environment and by contingencies in effect at other times. Expected reinforcer rates were calculated by entering simulated responding into a reinforcement feedback function. These were then fitted using Herrnstein's hyperbola, and the sampling distributions of the two fitted parameters were studied. Both k and Re were underestimated by curve fitting when low-deprivation or reinforcer-quality conditions were simulated. Further simulations showed that k and Re were increasingly underestimated as the assumed noise level was increased, particularly when low-deprivation or reinforcer quality was assumed. It is concluded that reported variations in k from single schedules should not be taken to indicate that the asymptotic rate of responding depends on reinforcement parameters.

Animals↗

Uncertainty analysis of parameters for modeling the transfer and fate of benzo(a)pyrene in Tianjin wastewater irrigated areas.

A Monte Carlo simulation for uncertainty analysis of three key parameters (local coal consumption rate Q(1L), dry deposition velocity of aerosol particulate Kp and biodegradation rate of benzo(a)pyrene in soil and sediment K(R3)) was conducted in this study. Results of the simulation indicate that the three parameters were influenced by uncertainty and that all equilibrium concentrations in the four bulk compartments and various sub-compartments were log-normally distributed. However, the results also indicated that among the six primary transfer fluxes, erosion associated with solids in soil and deposition associated with solids in water, along with output from sewers were also log-normally distributed, while deposition from air to soil and biodegradation in soil and sediment followed normal distributions. The effect of uncertainty on the model results of the three key parameters was derived using a comparison of upper and lower of confidence interval boundaries at the 95% level of confidence. The results reveal that uncertainty in the key parameters had a more significant influence on equilibrium concentrations of the chemical in the bulk compartments of soil and sediment than on concentrations in the other two bulk compartments, various sub-compartments and the six predominant transfer fluxes.

Benzo(a)pyrene↗

Bayesian restoration of a hidden Markov chain with applications to DNA sequencing.

Hidden Markov models (HMMs) are a class of stochastic models that have proven to be powerful tools for the analysis of molecular sequence data. A hidden Markov model can be viewed as a black box that generates sequences of observations. The unobservable internal state of the box is stochastic and is determined by a finite state Markov chain. The observable output is stochastic with distribution determined by the state of the hidden Markov chain. We present a Bayesian solution to the problem of restoring the sequence of states visited by the hidden Markov chain from a given sequence of observed outputs. Our approach is based on a Monte Carlo Markov chain algorithm that allows us to draw samples from the full posterior distribution of the hidden Markov chain paths. The problem of estimating the probability of individual paths and the associated Monte Carlo error of these estimates is addressed. The method is illustrated by considering a problem of DNA sequence multiple alignment. The special structure for the hidden Markov model used in the sequence alignment problem is considered in detail. In conclusion, we discuss certain interesting aspects of biological sequence alignments that become accessible through the Bayesian approach to HMM restoration.

Algorithms↗

Diurnal variation in supercooling points of three species of Collembola from Cape Hallett, Antarctica.

Daily changes in microclimate temperature and supercooling point (SCP) of Collembola were measured during summer at Cape Hallett, North Victoria Land, Antarctica. Isotoma klovstadi and Cryptopygus cisantarcticus (Isotomidae) showed bimodal SCP distributions, predominantly in the high group during the day and in the low group during the night. There were no concurrent diurnal changes in water content or haemolymph osmolality. By contrast, Friesea grisea (Neanuridae) had a unimodal distribution of SCPs that was invariant between daytime and nighttime. Isotoma klovstadi collected foraging on moss had uniformly high SCPs, which shifted towards the low group when the animals were starved for 2-8 h. When I. klovstadi was acclimated for five days with lichen or algae, SCPs were higher than if they were supplied with moss, while those that were starved (with free water or 100% relative humidity) displayed a trimodal SCP distribution. A variety of pre-treatments, including cold, heat, desiccation and slow cooling were ineffective at inducing SCP shifts in C. cisantarcticus or I. klovstadi. It is postulated that behavioural avoidance of low temperatures by vertical migration may be key in I. klovstadi's short-term survival of nighttime temperatures. These data suggest that the full range of thermal responses of Antarctic Collembola is yet to be elucidated.

Acclimatization↗

An E-M algorithm and testing strategy for multiple-locus haplotypes.

This paper gives an expectation maximization (EM) algorithm to obtain allele frequencies, haplotype frequencies, and gametic disequilibrium coefficients for multiple-locus systems. It permits high polymorphism and null alleles at all loci. This approach effectively deals with the primary estimation problems associated with such systems; that is, there is not a one-to-one correspondence between phenotypic and genotypic categories, and sample sizes tend to be much smaller than the number of phenotypic categories. The EM method provides maximum-likelihood estimates and therefore allows hypothesis tests using likelihood ratio statistics that have chi 2 distributions with large sample sizes. We also suggest a data resampling approach to estimate test statistic sampling distributions. The resampling approach is more computer intensive, but it is applicable to all sample sizes. A strategy to test hypotheses about aggregate groups of gametic disequilibrium coefficients is recommended. This strategy minimizes the number of necessary hypothesis tests while at the same time describing the structure of disequilibrium. These methods are applied to three unlinked dinucleotide repeat loci in Navajo Indians and to three linked HLA loci in Gila River (Pima) Indians. The likelihood functions of both data sets are shown to be maximized by the EM estimates, and the testing strategy provides a useful description of the structure of gametic disequilibrium. Following these applications, a number of simulation experiments are performed to test how well the likelihood-ratio statistic distributions are approximated by chi 2 distributions. In most circumstances the chi 2 grossly underestimated the probability of type I errors. However, at times they also overestimated the type 1 error probability. Accordingly, we recommended hypothesis tests that use the resampling method.

Algorithms↗

The Gompertz function does not measure ageing.

The Gompertz transform of the distribution function for the age at death expresses mortality in a form R = R0e(alphat) where R0 is the mortality at time zero and alpha is the rate of increase of mortality, frequently taken as the rate of ageing. The slope of the line alpha is frequently used as a measure of the rate of ageing. It is argued that it is incorrect to use alpha in this way. To support this contention, a paradox is produced whereby selection for longevity increases alpha, which could lead to the absurd conclusion that selection for longevity increases the rate of ageing.

Aging↗