Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “statistical inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 937 records · Page 52Linked to original sources

The relevance of drug injectors' social and risk networks for understanding and preventing HIV infection.

Focusing on the social environment as well as the individual should both enhance our understanding of HIV transmission and assist in the development of more effective prevention programs. Networks are an important aspect of drug injectors' social environment. We distinguish between (1) risk networks (the people among whom HIV risk behaviors occur) as vectors of disease transmission, and (2) social networks (the people among whom there are social interactions with a mutual orientation to one another) as generators and disseminators of social influence. These concepts are applied to analyses of data from interviews with drug injectors in two studies. In the first study drug injectors' risk networks converge with their social networks: 70% inject or share syringes with a spouse or sex partner, a running partner, or with friends or others whom they know. Qualitative data from interviews with injectors in the second study also show that the social relationships between drug injectors and members of their risk network are often based on long-standing and multiplex relationships, such as those based on kinship, friendship, marital and sexual ties, and economic activity. In the first study the vast majority of injectors, over 90%, have social ties with non-injectors. Injectors with more frequent social contacts with non-injectors engage in lower levels of injecting risk behavior. Risk settings may function as risk networks: injectors in this study who inject at shooting galleries are more likely than those who do not to rent used syringes, borrow used syringes and inject with strangers. Since the adoption of a network approach is relatively new, a number of issues require further attention. These include: how to utilize social networks among drug injectors to reduce risk through peer pressure; how to promote risk reduction by encouraging ties between injectors and non-injectors; and how to integrate biographical and historical change into understanding network processes. Appropriate methodologies to study drug injectors' networks should be developed, including techniques to reach hidden populations, computer software for managing and analyzing network data bases, and statistical methods for drawing inferences from data gathered through dependent sampling designs.

Adult↗

Molecular phylogeny of Eubacteria: a new multiple tree analysis method applied to 15 sequence data sets questions the monophyly of gram-positive bacteria.

Phylogenetic relationships between the major eubacterial phyla were studied using the sequences of 15 homologous bacterial genes. Neither the classical concatenation strategy nor a new multiple tree analysis method involving statistical tests of the inferred phylogenetic relationships provided any (solid) conclusions about eubacterial phylogeny; no pairs of eubacterial phyla proved to be closer to each other in the 15 reconstructed trees than would be expected for trees with random topologies. The phylogeny of Eubacteria therefore appears to be tightly bush-like. Moreover, results from both concatenation and multiple tree analysis raise doubts concerning the monophyly of the so-called Gram-positive bacteria phylum, since the monophyly hypothesis is no more strongly supported by data than its alternatives. It is noteworthy that the structural bases for the Gram-positive phenotype are not incompatible with the hypothesis of independent emergence of this character at two different times.

Bacteria↗

Discrepancies in cancer mortality estimates.

Progress against cancer-associated mortality could be estimated by mortality rates. However, population-based measures may not be comparable to trends across time based on individual cases followed over time. For two leading cancer types (colorectal and lung and bronchus), the following were calculated from the Surveillance, Epidemiology, and End Results (SEER) Program survival matrix output: probability of death (PD) as the relative cumulative deaths in a cohort of cancer patients at 12-month intervals after the year of diagnosis, 5-year survival probability, and median survival time. Annual age-adjusted U.S. mortality rates were obtained from the National Center for Health Statistics (NCHS). Colorectal cancer PD 5-year (patients having survived up to 5 years after diagnosis) decreased 4.3% from 1985-1997, in contrast to a 20% decrease in the mortality rate during the same period. The mortality rate for lung and bronchus cancer sharply increased between 1973 and 1991 and was followed by a clear downward trend of 2.5% from 1991-1997, while its PD 5-year decreased by only 0.79%. PD for specific cohorts at yearly intervals did not improve across time at comparable intervals; however, the mortality rate was clearly reduced for both anatomic sites. Survival and PD were, as expected, inversely related. Changes in median survival and 5-year survival probability were similar over time. Although the cancer mortality rate is a clear statistical end-point, inference of success by its falling trend alone could be unrelated to the SEER probability of death.

Humans↗

Removal of boron from aqueous solution by adsorption on Al2O3 based materials using full factorial design.

This paper aims the adsorption of boron from aqueous solution onto Siral 30 and Pural using 2(3) full factorial design. The effect of individual variables and their interactional effects for boron adsorption were also determined. From the statistical analysis, it is inferred that as pH and temperature increased boron adsorption from aqueous solution decreased. Siral 30 was found to be more efficient adsorbent than Pural. The unimportant factor affecting boron adsorption from aqueous solution was also verified by using Fisher adequacy test. At the 90% confidence level, the type of adsorbent, temperature and type of adsorbent-temperature interaction was effective on boron adsorption from aqueous solution. The experimental results were fitted to the Langmuir, Freundlich and Dubinin-Radushkevich (DR) equations to find out adsorption capacities. In most cases, the results indicate that Freundlich and DR equations are well described with the sorption data. The adsorption capacity values of Siral 30 calculated from Freundlich and DR equation was greater than that of Pural. The thermodynamic parameters were also estimated and the adsorption process was not spontaneous nature.

Adsorption↗

Bootstrapped DEPICT for error estimation in PET functional imaging.

Basis pursuit denoising is a new approach for data-driven estimation of parametric images from dynamic positron emission tomography (PET) data. At present, this kinetic modeling technique does not allow for the estimation of the errors on the parameters. These estimates are useful when performing subsequent statistical analysis, such as, inference across a group of subjects or when applying partial volume correction algorithms. The difficulty with calculating the error estimates is a consequence of using an overcomplete dictionary of kinetic basis functions. In this paper, a bootstrap approach for the estimation of parameter errors from dynamic PET data is presented. This paper shows that the bootstrap can be used successfully to compute parameter errors on a region of interest or parametric image basis. Validation studies evaluate the methods performance on simulated and measured PET data ([(11)C]Diprenorphine-opiate receptor and [(11)C]Raclopride-dopamine D(2) receptor). The method is presented in the context of PET neuroreceptor binding studies, however, it has general applicability to a wide range of PET/SPET radiotracers in neurology, oncology and cardiology.

Adult↗

Estimation of genotype error rate using samples with pedigree information--an application on the GeneChip Mapping 10K array.

Currently, most analytical methods assume all observed genotypes are correct; however, it is clear that errors may reduce statistical power or bias inference in genetic studies. We propose procedures for estimating error rate in genetic analysis and apply them to study the GeneChip Mapping 10K array, which is a technology that has recently become available and allows researchers to survey over 10,000 SNPs in a single assay. We employed a strategy to estimate the genotype error rate in pedigree data. First, the "dose-response" reference curve between error rate and the observable error number were derived by simulation, conditional on given pedigree structures and genotypes. Second, the error rate was estimated by calibrating the number of observed errors in real data to the reference curve. We evaluated the performance of this method by simulation study and applied it to a data set of 30 pedigrees genotyped using the GeneChip Mapping 10K array. This method performed favorably in all scenarios we surveyed. The dose-response reference curve was monotone and almost linear with a large slope. The method was able to estimate accurately the error rate under various pedigree structures and error models and under heterogeneous error rates. Using this method, we found that the average genotyping error rate of the GeneChip Mapping 10K array was about 0.1%. Our method provides a quick and unbiased solution to address the genotype error rate in pedigree data. It behaves well in a wide range of settings and can be easily applied in other genetic projects. The robust estimation of genotyping error rate allows us to estimate power and sample size and conduct unbiased genetic tests. The GeneChip Mapping 10K array has a low overall error rate, which is consistent with the results obtained from alternative genotyping assays.

Computer Simulation↗

Widespread anatomical abnormalities of grey and white matter structure in tuberous sclerosis.

BACKGROUND: Neuroimaging studies of tuberous sclerosis complex (TSC) have previously focused mainly on tubers or subependymal nodules. Subtle pathological changes in the structure of the brain have not been studied in detail. Computationally intensive techniques for reliable morphometry of brain structure are useful in disorders like TSC, where there is little prior data to guide selection of regions of interest. METHODS: Dual-echo, fast spin-echo MRI data were acquired from 10 TSC patients of normal intelligence and eight age-matched controls. Between-group differences in grey matter, white matter and cerebrospinal fluid were estimated at each intracerebral voxel after registration of these images in standard space; a permutation test based on spatial statistics was used for inference. CSF-attenuated FLAIR images were acquired for neuroradiological rating of tuber number. RESULTS: Significant deficits were found in patients, relative to comparison subjects, of grey matter volume bilaterally in the medial temporal lobes, posterior cingulate gyrus, thalamus and basal ganglia, and unilaterally in right fronto-parietal cortex (patients -20%). We also found significant and approximately symmetrical deficits of central white matter involving the longitudinal fasciculi and other major intrahemispheric tracts (patients -21%); and a bilateral cerebellar region of relative white matter excess (patients +28%). Within the patient group, grey matter volume in limbic and subcortical regions of deficit was negatively correlated with tuber count. CONCLUSIONS: Neuropathological changes associated with TSC may be more extensive than hitherto suspected, involving radiologically normal parenchymal structures as well as tubers, although these two aspects of the disorder may be correlated.

Adult↗

Computational models for the helix tilt angle.

The concept of hydrophobic imbalance and that of hydrophobic and hydrophilic centers are used along with side chain models in the computation of helix orientation and tilt angle in or near a membrane. Rotamer statistics are used to infer typical side chain positions and chain length for each amino acid, and the results are used in fast computation of helix orientation. Sliding windows are used to compute local tilt angles on long alpha-helices that defy idealized modeling and generate tilt angle profiles. Seven different procedures based on different formulas and hydrophobicity scales are used for comparison. These procedures generated very similar tilt angle profiles. These profiles provide insights into helix deformation, membrane destabilization, and similarity and differences between membrane proteins.

Algorithms↗

Bayesian analysis of population PK/PD models: general concepts and software.

Markov chain Monte Carlo (MCMC) techniques have revolutionized the field of Bayesian statistics by enabling posterior inference for arbitrarily complex models. The now widely used WinBUGS software has, over the years, made the methodology accessible to a great many applied scientists, in all fields of research. Despite this, serious application of MCMC methods within the field of population PK/PD has been comparatively limited. We appreciate that for many applied pharmacokineticists the prospect of conducting a Bayesian analysis will require numerous alien concepts to be taken on board and it may be difficult to justify investing the time and effort required in order to understand them (especially since the approach is so computer-intensive). For this reason we provide here a thorough (but often informal) discussion of all aspects of Bayesian inference as they apply specifically to population PK/PD. We also acknowledge that while the WinBUGS software is general purpose, model specification for some types of problem, population PK/PD being a prime example, can be very difficult, to the extent that a specialized interface for describing the problem at hand is often a practical necessity. In the latter part of this paper we describe such an interface, namely PKBugs. A principal aim of the paper is to offer sufficient technical background, in an easy to follow format, that the reader may develop both the confidence and know-how to make appropriate use of the PKBugs/WinBUGS framework (or similar software) for their own data analysis needs, should they choose to adopt a Bayesian approach.

Bayes Theorem↗

The distribution of the ancestral haplotype in finite stepping-stone models with population expansion.

Recent extensive analyses of human DNA polymorphism reveal that the ancestral haplotype at various genetic loci occurs almost exclusively in African samples. We develop a coalescence-based simulation method in stepping-stone models with population expansion and examine the probability (P(A)) that the ancestral haplotype is found in African samples and the probability (Q(A)) that the most recent common ancestor of sampled genes occurs in Africa. These probabilities and other summary statistics are used to infer the human demographic history. It is shown that the high observed P(A) value cannot be explained simply by sampling bias. Rather, it suggests that the African population has been more strongly subdivided and isolated from each other than the non-African population and that there must have been some African populations which were not directly involved in the Out-of-Africa expansion in the late Pleistocene.

Africa↗

Human immunodeficiency virus type 1 molecular evolution and the measure of selection.

Human immunodeficiency virus (HIV) envelope genes are highly variable between and often within individuals. Part of this variability is thought to be the result of immune-mediated positive selection for sequence diversity. To measure positive selection it has become customary in HIV research to calculate the ratio of the proportions of synonymous (ds) and nonsynonymous (dn) substitutions per potential synonymous or nonsynonymous site, respectively. However, another measure that can be used is the difference between ds and dn, delta d. We show, by example, that using the ratio, ds/dn, or the difference, delta d, may lead us to different conclusions regarding the existence of positive selection pressure. We conclude by noting that until we understand the processes that mediate nucleotide variation in a host selective environment, inferences based on summary statistics characterizing types of nucleotide substitutions should be made with caution.

Base Sequence↗

Gene selection and clustering for time-course and dose-response microarray experiments using order-restricted inference.

We propose an algorithm for selecting and clustering genes according to their time-course or dose-response profiles using gene expression data. The proposed algorithm is based on the order-restricted inference methodology developed in statistics. We describe the methodology for time-course experiments although it is applicable to any ordered set of treatments. Candidate temporal profiles are defined in terms of inequalities among mean expression levels at the time points. The proposed algorithm selects genes when they meet a bootstrap-based criterion for statistical significance and assigns each selected gene to the best fitting candidate profile. We illustrate the methodology using data from a cDNA microarray experiment in which a breast cancer cell line was stimulated with estrogen for different time intervals. In this example, our method was able to identify several biologically interesting genes that previous analyses failed to reveal.

Algorithms↗

Confirmation of data mining based predictions of protein function.

MOTIVATION: A central problem in bioinformatics is the assignment of function to sequenced open reading frames (ORFs). The most common approach is based on inferred homology using a statistically based sequence similarity (SIM) method, e.g. PSI-BLAST. Alternative non-SIM based bioinformatic methods are becoming popular. One such method is Data Mining Prediction (DMP). This is based on combining evidence from amino-acid attributes, predicted structure and phylogenic patterns; and uses a combination of Inductive Logic Programming data mining, and decision trees to produce prediction rules for functional class. DMP predictions are more general than is possible using homology. In 2000/1, DMP was used to make public predictions of the function of 1309 Escherichia coli ORFs. Since then biological knowledge has advanced allowing us to test our predictions. RESULTS: We examined the updated (20.02.02) Riley group genome annotation, and examined the scientific literature for direct experimental derivations of ORF function. Both tests confirmed the DMP predictions. Accuracy varied between rules, and with the detail of prediction, but they were generally significantly better than random. For voting rules, accuracies of 75-100% were obtained. Twenty-one of these DMP predictions have been confirmed by direct experimentation. The DMP rules also have interesting biological explanations. DMP is, to the best of our knowledge, the first non-SIM based prediction method to have been tested directly on new data. AVAILABILITY: We have designed the "Genepredictions" database for protein functional predictions. This is intended to act as an open repository for predictions for any organism and can be accessed at http://www.genepredictions.org

Abstracting and Indexing↗

2SNP: scalable phasing based on 2-SNP haplotypes.

2SNP software package implements a new very fast scalable algorithm for haplotype inference based on genotype statistics collected only for pairs of SNPs. This software can be used for comparatively accurate phasing of large number of long genome sequences, e.g. obtained from DNA arrays. As an input 2SNP takes genotype matrix and outputs the corresponding haplotype matrix. On datasets across 79 regions from HapMap 2SNP is several orders of magnitude faster than GERBIL and PHASE while matching them in quality measured by the number of correctly phased genotypes, single-site and switching errors. For example, 2SNP requires 41 s on Pentium 4 2 Ghz processor to phase 30 genotypes with 1381 SNPs (ENm010.7p15:2 data from HapMap) versus GERBIL and PHASE requiring more than a week and admitting no less errors than 2SNP.

Algorithms↗

Popper's philosophy for epidemiologists.

This paper discusses the application of Popper's philosophy to epidemiological research, examining in particular the problems of replication without risk of refutation, of mistaking statistical sophistication for deductive inference, and of dealing with causality at a general level. An example is given of a Popperian approach to the test of a causal hypothesis concerning cancer of the cervix.

Adolescent↗

Molecular haplotyping at high throughput.

Reconstruction of haplotypes, or the allelic phase, of single nucleotide polymorphisms (SNPs) is a key component of studies aimed at the identification and dissection of genetic factors involved in complex genetic traits. In humans, this often involves investigation of SNPs in case/control or other cohorts in which the haplotypes can only be partially inferred from genotypes by statistical approaches with resulting loss of power. Moreover, alternative statistical methodologies can lead to different evaluations of the most probable haplotypes present, and different haplotype frequency estimates when data are ambiguous. Given the cost and complexity of SNP studies, a robust and easy-to-use molecular technique that allows haplotypes to be determined directly from individual DNA samples would have wide applicability. Here, we present a reliable, automated and high-throughput method for molecular haplotyping in 2 kb, and potentially longer, sequence segments that is based on the physical determination of the phase of SNP alleles on either of the individual paternal haploids. We demonstrate that molecular haplotyping with this technique is not more complicated than SNP genotyping when implemented by matrix-assisted laser desorption/ionisation mass spectrometry, and we also show that the method can be applied using other DNA variation detection platforms. Molecular haplotyping is illustrated on the well-described beta(2)-adrenergic receptor gene.

Alleles↗

Evaluation of methods for detecting recombination from DNA sequences: empirical data.

The performance of 14 different recombination detection methods was evaluated by analyzing several empirical data sets where the presence of recombination has been suggested or where recombination is assumed to be absent. In general, recombination methods seem to be more powerful with increasing levels of divergence, but different methods showed distinct performance. Substitution methods using summary statistics gave more accurate inferences than most phylogenetic methods. However, definitive conclusions about the presence of recombination should not be derived on the basis of a single method. Performance patterns observed from the analysis of real data sets coincided very well with previous computer simulation results. Previous recombination inferences from some of the data sets analyzed here should be reconsidered. In particular, recombination in HIV-1 seems to be much more widespread than previously thought. This finding might have serious implications on vaccine development and on the reliability of previous inferences of HIV-1 evolutionary history and dynamics.

DNA↗

Neural network applications in physical medicine and rehabilitation.

The purpose of this article is to provide an overview of neural networks and their applications in physical medicine and rehabilitation. Conventional statistical models may present certain limitations that can be overcome by neural networks. We show what neural networks are, how they "learn" regularities from the data, and how they can classify previously unseen cases. We present advantages and disadvantages of using neural networks and compare them with regression models. We explain how neural networks can be used as statistical tools for making inferences using the example of a prognostic model that predicts ambulation after spinal cord injury.

Humans↗