Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “statistical inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 829 records · Page 46Linked to original sources

Analysis of antimalarial synergy between bestatin and endoprotease inhibitors using statistical response-surface modelling.

The pathway of hemoglobin degradation by erythrocytic stages of the human malarial parasite Plasmodium falciparum involves initial cleavages of globin chains, catalyzed by several endoproteases, followed by liberation of amino acids from the resulting peptides, probably by aminopeptidases. This pathway is considered a promising chemotherapeutic target, especially in view of the antimalarial synergy observed between inhibitors of aspartyl and cysteine endoproteases. We have applied response-surface modelling to assess antimalarial interactions between endoprotease and aminopeptidase inhibitors using cultured P. falciparum parasites. The synergies observed were consistent with a combined role of endoproteases and aminopeptidases in hemoglobin catabolism in this organism. As synergies between antimicrobial agents are often inferred without proper statistical analysis, the model used may be widely applied in studies of antimicrobial drug interactions.

Algorithms↗

Allele excess at neutrally evolving microsatellites and the implications for tests of neutrality.

Skews in the observed allele-frequency spectrum are frequently viewed as an indication of non-neutral evolution. Recent surveys of microsatellite variability have used an excess of alleles as a statistical approach to infer positive selection. Using neutral coalescent simulations we demonstrate that the mean numbers of alleles expected under the stepwise-mutation model and infinite-allele model deviate from the observed numbers of alleles. The magnitude of this difference is dependent on the sample size, mutation rates (theta-values) and observed gene diversities. Moreover, we show that the number of observed alleles differs among loci with the same observed gene diversity but different mutation rates (theta-values). We propose that a reliable test statistic based on allele excess must determine the confidence interval by computer simulations conditional on the observed gene diversity and theta-values. As the latter are notoriously difficult to obtain for experimental data, we suggest that other statistics, such as lnRV, may be better suited to the identification of microsatellite loci subject to selection.

Biological Evolution↗

Extensive gene flow blurs phylogeographic but not phylogenetic signal in Olea europaea L.

Genetic structure and evolutionary patterns of the wild olive tree (Olea europaea L.) were investigated with AFLP fingerprinting data at three geographic levels: (a) phylogenetic relationships of the six currently recognized subspecies in Eurasia and Africa; (b) lineage identification in subsp. europaea of the Mediterranean basin; and (c) phylogeography in the western Mediterranean. Two statistical approaches (Bayesian inference and analysis of molecular variance) were used to analyse the AFLP fingerprints. To determine the congruency and transferability of results across studies previous RAPD and ISSR data were analysed in a similar manner. Comparisons proved that qualitative results were mostly congruent but quantitative values differed, depending on the method of analysis. Neighbour-Joining analysis of AFLP phenotypes supported current classification of subspecies. At a Mediterranean scale no clear cut phylogeographic pattern was recovered, likely due to extensive gene flow between populations of subsp. europaea. Gene flow estimates calculated with conventional F-statistics showed that reproductive barriers separated neither populations nor lineages of O. europaea. Genetic divergence between eastern and western parts of the Mediterranean basin was observed only when geographical and population information were incorporated into the analyses through hierarchical analysis of molecular variance (AMOVA). Within the western Mediterranean, the highest genetic diversity was found in two regions: on both sides of the Strait of Gibraltar and in the Balearic archipelago. Additionally, long-lasting isolation of the northern-most populations of the Iberian Peninsula appeared to be responsible for a significant divergence.

DNA Fingerprinting↗

GENIE: estimating demographic history from molecular phylogenies.

UNLABELLED: GENIE implements a statistical framework for inferring the demographic history of a population from phylogenies that have been reconstructed from sampled DNA sequences. The methods are based on population genetic models known collectively as coalescent theory. AVAILABILITY: GENIE is available from http://evolve.zoo.ox.ac.uk. All popular operating systems are supported.

Computer Simulation↗

Applications of the Mantel-Haenszel statistic to the comparison of survival distributions.

In comparing two survival distributions, a Mantel-Haenszel statistic can be computed after each death as a non-linear two-sample rank statistic. The distributions of both the maximum and terminal statistics in such a sequence are studied numerically, in the absence of censoring, and appropriate critical values are determined. The maximum statistic is applied to simultaneous inference, and both the maximum and terminal statistics are used as the basis for early stopping procedures (especially in the pseudo-sequential context). Procedures based on the two statistics are compared for power and for early decision properties such as stopping index and (for exponential distributions) stopping time.

Biometry↗

Template-based recognition of protein fold within the midnight and twilight zones of protein sequence similarity.

Most homologous pairs of proteins have no significant sequence similarity to each other and are not identified by direct sequence comparison or profile-based strategies. However, multiple sequence alignments of low similarity homologues typically reveal a limited number of positions that are well conserved despite diversity of function. It may be inferred that conservation at most of these positions is the result of the importance of the contribution of these amino acids to the folding and stability of the protein. As such, these amino acids and their relative positions may define a structural signature. We demonstrate that extraction of this fold template provides the basis for the sequence database to be searched for patterns consistent with the fold, enabling identification of homologs that are not recognized by global sequence analysis. The fold template method was developed to address the need for a tool that could comprehensively search the midnight and twilight zones of protein sequence similarity without reliance on global statistical significance. Manual implementations of the fold template method were performed on three folds--immunoglobulin, c-lectin and TIM barrel. Following proof of concept of the template method, an automated version of the approach was developed. This automated fold template method was used to develop fold templates for 10 of the more populated folds in the SCOP database. The fold template method developed three-dimensional structural motifs or signatures that were able to return a diverse collection of proteins, while maintaining a low false positive rate. Although the results of the manual fold template method were more comprehensive than the automated fold template method, the diversity of the results from the automated fold template method surpassed those of current methods that rely on statistical significance to infer evolutionary relationships among divergent proteins.

Amino Acid Sequence↗

Representation and timing in theories of the dopamine system.

Although the responses of dopamine neurons in the primate midbrain are well characterized as carrying a temporal difference (TD) error signal for reward prediction, existing theories do not offer a credible account of how the brain keeps track of past sensory events that may be relevant to predicting future reward. Empirically, these shortcomings of previous theories are particularly evident in their account of experiments in which animals were exposed to variation in the timing of events. The original theories mispredicted the results of such experiments due to their use of a representational device called a tapped delay line. Here we propose that a richer understanding of history representation and a better account of these experiments can be given by considering TD algorithms for a formal setting that incorporates two features not originally considered in theories of the dopaminergic response: partial observability (a distinction between the animal's sensory experience and the true underlying state of the world) and semi-Markov dynamics (an explicit account of variation in the intervals between events). The new theory situates the dopaminergic system in a richer functional and anatomical context, since it assumes (in accord with recent computational theories of cortex) that problems of partial observability and stimulus history are solved in sensory cortex using statistical modeling and inference and that the TD system predicts reward using the results of this inference rather than raw sensory data. It also accounts for a range of experimental data, including the experiments involving programmed temporal variability and other previously unmodeled dopaminergic response phenomena, which we suggest are related to subjective noise in animals' interval timing. Finally, it offers new experimental predictions and a rich theoretical framework for designing future experiments.

Algorithms↗

Natural selection and the molecular clock.

This paper concludes that the statistical properties of protein evolution are compatible with a particular model of evolution by natural selection. The argument begins with a statistical description of the molecular clock based on a Poisson process with a randomly varying tick rate. If the time scale of the change of the tick rate of the molecular clock is assumed to be much less than the average time between substitutions, then it is shown that the substitution process must be episodic, with bursts of substitutions being separated by long periods of time with no substitutions. This analysis generalizes the recent work of Gillespie (1984a). The second part of the argument shows that a simple model of evolution by natural selection--one that incorporates a changing environment, the molecular landscape, and a simple form of epistasis--exhibits dynamics that are identical to those inferred from the statistical analysis. This leads to the conclusion that natural selection is a viable explanation for protein evolution. In addition, a correction formula for multiple substitutions is given that does not require that the substitution process be a Poisson process, and some comments on the inability of the neutral allele theory to account for the dynamics of the substitution process are presented.

Animals↗

What they want and what they get: the social goals of boys with ADHD and comparison boys.

Twenty-seven boys diagnosed with attention-deficit hyperactivity disorder (ADHD) and 18 comparison boys participated in a competitive tetradic interaction task. Boys were individually interviewed before the game about their goals for the interaction, and adult observers inferred boys' social goals from videotapes of the interaction. Social acceptance was determined by combining positive and negative sociometric nominations collected through individual interviews at the end of the summer research program in which the interaction was held. In their self-reports, ADHD-high aggressive boys prioritized trouble-seeking and fun at the expense of rules to a greater extent than did both ADHD-low aggressive and comparison boys. Observers judged ADHD-high aggressive boys to seek attention more strongly and seek fairness less strongly than of the other two groups. Self-reported goals of defiance and cooperation predicted boys' end-of-program social standing, even with interactional behaviors and subgroup status controlled statistically. Observer-inferred goals were differentially associated with social acceptance for ADHD and comparison boys, suggesting discontinuities in peer interaction processes. Differentiation of goals from behavior and the integral role of children's goals in peer acceptance are discussed.

Aggression↗

Phylogeny and evolution of the major intrinsic protein family.

BACKGROUND INFORMATION: MIPs (major intrinsic proteins) form channels across biological membranes that control recruitment of water and small solutes such as glycerol and urea in all living organisms. Because of their widespread occurrence and large number, MIPs are a sound model system to understand evolutionary mechanisms underlying the generation of protein structural and functional diversity. With the recent increase in genomic projects, there is a considerable increase in the quantity and taxonomic range of MIPs in molecular databases. RESULTS: In the present study, I compiled more than 450 non-redundant amino acid sequences of MIPs from NCBI databases. Phylogenetic analyses using Bayesian inference reconstructed a statistically robust tree that allowed the classification of members of the family into two main evolutionary groups, the GLPs (glycerol-uptake facilitators or aquaglyceroporins) and the water transport channels or AQPs (aquaporins). Separate phylogenetic analyses of each of the MIP subfamilies were performed to determine the main groups of orthology. In addition, comparative sequence analyses were conducted to identify conserved signatures in the MIP molecule. CONCLUSIONS: The earliest and major gene duplication event in the history of the MIP family led to its main functional split into GLPs and AQPs. GLPs show typically one single copy in microbes (eubacteria, archaea and fungi), up to four paralogues in vertebrates and they are absent from plants. AQPs are usually single in microbes and show their greatest numbers and diversity in angiosperms and vertebrates. Functional recruitment of NOD26-like intrinsic proteins to glycerol transport due to the absence of GLPs in plants was highly supported. Acquisition of other MIP functions such as permeability to ammonia, arsenite or CO2 is restricted to particular MIP paralogues. Up to eight fairly conserved boxes were inferred in the primary sequence of the MIP molecule. All of them mapped on to one side of the channel except the conserved glycine residues from helices 2 and 5 that were found in the opposite side.

Amino Acid Sequence↗

Primary-consistent soft-decision color demosaicking for digital cameras (patent pending).

Color mosaic sampling schemes are widely used in digital cameras. Given the resolution of CCD sensor arrays, the image quality of digital cameras using mosaic sampling largely depends on the performance of the color demosaicking process. A common problem with existing color demosaicking algorithms is an inconsistency of sample interpolations in different primary color channels, which is the cause of the most objectionable color artifacts. To cure the problem, we propose a new primary-consistent soft-decision framework (PCSD) of color demosaicking. In the PCSD framework, we make multiple estimates of a missing color sample under different hypotheses on edge or texture directions. The estimates are made via a primary consistent interpolation, meaning that all three primary components of a color are interpolated in the same direction. The final estimate of a color sample is obtained by testing different interpolation hypotheses in the reconstructed full-resolution color image and selecting the best via an optimal statistical decision or inference process. A concrete color demosaicking method of the PCSD framework is presented. This new method eliminates certain types of color artifacts of existing color demosaicking methods. Extensive experimental results demonstrate that the PCSD approach can significantly improve the image quality of digital cameras in both subjective and objective measures. In some instances, our gain over the competing methods can be as much as 7 dB.

Algorithms↗

To smooth or not to smooth? Bias and efficiency in fMRI time-series analysis.

This paper concerns temporal filtering in fMRI time-series analysis. Whitening serially correlated data is the most efficient approach to parameter estimation. However, if there is a discrepancy between the assumed and the actual correlations, whitening can render the analysis exquisitely sensitive to bias when estimating the standard error of the ensuing parameter estimates. This bias, although not expressed in terms of the estimated responses, has profound effects on any statistic used for inference. The special constraints of fMRI analysis ensure that there will always be a misspecification of the assumed serial correlations. One resolution of this problem is to filter the data to minimize bias, while maintaining a reasonable degree of efficiency. In this paper we present expressions for efficiency (of parameter estimation) and bias (in estimating standard error) in terms of assumed and actual correlation structures in the context of the general linear model. We show that: (i) Whitening strategies can result in profound bias and are therefore probably precluded in parametric fMRI data analyses. (ii) Band-pass filtering, and implicitly smoothing, has an important role in protecting against inferential bias.

Algorithms↗

A statistical approach designed for finding mathematically defined repeats in shotgun data and determining the length distribution of clone-inserts.

The large amount of repeats, especially high copy repeats, in the genomes of higher animals and plants makes whole genome assembly (WGA) quite difficult. In order to solve this problem, we tried to identify repeats and mask them prior to assembly even at the stage of genome survey. It is known that repeats of different copy number have different probabilities of appearance in shotgun data, so based on this principle, we constructed a statistical model and inferred criteria for mathematically defined repeats (MDRs) at different shotgun coverages. According to these criteria, we developed software MDRmasker to identify and mask MDRs in shotgun data. With repeats masked prior to assembly, the speed of assembly was increased with lower error probability. In addition, clone-insert size affect the accuracy of repeat assembly and scaffold construction, we also designed length distribution of clone-inserts using our model. In our simulated genomes of human and rice, the length distribution of repeats is different, so their optimal length distributions of clone-inserts were not the same. Thus with optimal length distribution of clone-inserts, a given genome could be assembled better at lower coverage.

Animals↗

Recent developments in the analysis of comparative data.

Comparative methods can be used to test ideas about adaptation by identifying cases of either parallel or convergent evolutionary change across taxa. Phylogenetic relationships must be known or inferred if comparative methods are to separate the cross-taxonomic covariation among traits associated with evolutionary change from that attributable to common ancestry. Only the former can be used to test ideas linking convergent or parallel evolutionary change to some aspect of the environment. The comparative methods that are currently available differ in how they manage the effects brought about by phylogenetic relationships. One method is applicable only to discrete data, and uses cladistic techniques to identify evolutionary events that depart from phylogenetic trends. Techniques for continuous variables attempt to control for phylogenetic effects in a variety of ways. One method examines the taxonomic distribution of variance to identify the taxa within which character variation is small. The method assumes that taxa with small amounts of variation are those in which little evolutionary change has occurred, and thus variation is unlikely to be independent of ancestral trends. Analyses are then concentrated among taxa that show more variation, on the assumption that greater evolutionary change in the character has taken place. Several methods estimate directly the extent to which ancestry can predict the observed variation of a character, and subtract the ancestral effect to reveal variation of phylogeny. Yet another can remove phylogenetic effects if the true phylogeny is known. One class of comparative methods controls for phylogenetic effects by searching for comparative trends within rather than across taxa. With current knowledge of phylogenies, there is a trade-off in the choice of a comparative method: those that control phylogenetic effects with greater certainty are either less applicable to real data, or they make restrictive or untestable assumptions. Those that rely on statistical patterns to infer phylogenetic effects may not control phylogeny as efficiently but are more readily applied to existing data sets.

Adaptation, Physiological↗

MDL and the statistical mechanics of protein potentials.

The combination of a wealth of structural data and impressive computational power provides detailed information pertaining to the structure and dynamics of biomacromolecules. A natural inclination is to incorporate this information into models to gain added predictive power on protein folding and stability. There has been considerable recent interest in developing "knowledge-based" potentials to describe internal interactions in proteins. In these approaches, probability distribution functions are inferred from existing knowledge. A common assumption has been the "quasi-chemical approximation" or "Boltzmann device". This method relates statistical mechanical probabilities to observed frequencies. The validity of this approach is discussed in detail from a statistical mechanics perspective. Because statistical mechanics is a form of statistical inference based on a lack of knowledge of the system, the "Boltzmann device" does not have a rigorous theoretical justification. In the present work, a statistical mechanics based on partial knowledge of the system is employed. This statistical mechanical scheme uses the minimum description length (MDL) of phase space as its main tool. With this approach, "knowledge-based" potentials can be derived in a rigorous fashion. In practical calculations, these potentials are best obtained using Bayesian inference methods similar to those used in image reconstruction.

Algorithms↗

Predicting the conformational states of cyclic tetrapeptides.

Biologically active cyclic tetrapeptides, usually found among fungi metabolites, exhibit phytotoxic or cytostatic activities that are likely to be governed by specific conformations adopted in solution. For conformational studies and drug design, there is a strong interest in using fast and reliable methods to determine correctly the conformational population of cyclotetrapeptides. We show here that standard molecular mechanics computational approach gives satisfactory results. The method was validated step by step by experimental data either obtained after synthesis and NMR analysis, or found in the literature. The cyclo(Gly)(4), cyclo(Ala)(4), cyclo(Sar)(4), and cyclo(SarGly)(2) peptides were used to evaluate the prediction of the peptide backbone conformation, and the detailed conformational analysis of tentoxin, a natural phytotoxic cyclotetrapeptide in which N-alkylated peptide bonds alternate with regular secondary ones, was used to validate the computation of conformers proportions. From the knowledge of an initial cyclic primary structure and of the D or L configuration of the amino acids, we show that it is possible to determine the exact orientation of carbonyl groups and to predict the nature of conformers present in solution. The proportion of each conformer can be inferred from a statistical thermodynamics approach by using the potential energy values of each conformer, computed by molecular mechanics methods with the TRIPOS force field, which allowed us to account for the solvent. The solvent contribution was processed by two different methods according to the nature of the interactions: whether through the dielectric constant introduced in the electrostatic potential, when interaction with solute molecules are weak or negligible, or through the computation of free energy of solvation using the algorithm SILVERWARE for solvents explicitly interacting with the solute. When applied to tentoxin, this conformational analysis yielded results in very good agreement with the experimental data reported by Pinet et al. (Biopolymers, 1995, Vol. 36, pp. 135-152), on both the nature of existing conformers and their relative proportions, whatever the nature of the considered solvent.

Algorithms↗

Kinetics of biodegradation of binary and ternary mixtures of PAHs.

The kinetics of biodegradation of mixtures of polycyclic aromatic hydrocarbons (PAHs) by Sphingomonas paucimobilis strain EPA505 were investigated. The investigation focused on three- and four-ring PAHs, specifically 2-methylphenanthrene, fluoranthene, and pyrene. Uptake rates in aerobic batch suspended cultivations were measured for the individual PAHs and their binary and ternary mixtures. It was observed that kinetics were influenced by the mixture composition and the kinetic properties of the components. A material balance equation containing the Monod model was numerically fitted to uptake data to determine extant kinetic parameters for the individual PAHs. Similarly, equations containing kinetic interaction models derived from enzyme kinetics were fitted to the uptake data obtained from experiments with binary and ternary mixtures. The investigation considered the following interaction types: no-interaction (Monod), pure competitive interaction, noncompetitive or mixed-type interaction, uncompetitive inhibition, and nonspecific interaction based on pure competition (SKIP). Model fit was evaluated based on probabilistic and statistical criteria and inferences were reached about underlying interaction mechanisms based on model fit. Mixture kinetics were most adequately simulated by the pure competitive interaction model with mutual substrate exclusivity. This model is fully predictive, relying only on parameters determined in the sole-PAH experiments. It was shown that for low percent inhibition values and with limited data, pure competitive interaction kinetics may not be evident, resembling no-interaction kinetics. This study is a reasonable starting point for understanding and modeling biodegradation of complex PAH mixtures in engineered and natural systems.

Biodegradation, Environmental↗

Bayesian technique for investigating linearity in event-related BOLD fMRI.

Event-related BOLD fMRI data is modeled as a linear time-invariant system. Together with Bayesian inference techniques, a statistical test is developed for rigorously detecting linearity/nonlinearity in the BOLD response system. The test is applied to data collected from eight subjects using an event-related paradigm with a switching checkerboard as the visual stimulus. Analyzed as a group, the results clearly find the response to be nonlinear. When each subject is analyzed individually, however, the results are predominantly nonlinear, but there is some evidence to suggest that there may be a crossover from a linear to a nonlinear regime and vice versa. This could be important when estimating physiological parameters for individuals. Additionally, estimates of the hemodynamic response function and corresponding response were obtained, but there was no consistent appearance of a poststimulus undershoot in the event-related BOLD response.

Adult↗