Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian modelling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 973 records · Page 54Linked to original sources

A probabilistic approach to interpreting verbal autopsies: methodology and preliminary validation in Vietnam.

AIMS: Verbal autopsy (VA) has become an important tool in the past 20 years for determining cause of death in communities where there is no routine registration. In many cases, expert physicians have been used to interpret the VA findings and so assign individual causes of death. However, this is time consuming and not always repeatable. Other approaches such as algorithms and neural networks have been developed in some settings. This paper aims to develop a method that is simple, reliable and consistent, which could represent an advance in VA interpretation. METHODS: This paper describes the development of a Bayesian probability model for VA interpretation as an attempt to find a better approach. This methodology and a preliminary implementation are described, with an evaluation based on VA material from rural Vietnam. RESULTS: The new model was tested against a series of 189 VA interviews from a rural community in Vietnam. Using this very basic model, over 70% of individual causes of death corresponded with those determined by two physicians increasing to over 80% if those cases ascribed to old age or as being indeterminate by the physicians were excluded. DISCUSSION: Although there is a clear need to improve the preliminary model and to test more extensively with larger and more varied datasets, these preliminary results suggest that there may be good potential in this probabilistic approach.

Autopsy↗

Some statistical and regulatory issues in the evaluation of genetic and genomic tests.

The genomics revolution is reverberating throughout the worlds of pharmaceutical drugs, genetic testing and statistical science. This revolution, which uses single nucleotide polymorphisms (SNPs) and gene expression technology, including cDNA and oligonucleotide microarrays, for a range of tests from home-brews to high-complexity lab kits, can allow the selection or exclusion of patients for therapy (responders or poor metabolizers). The wide variety of US regulatory mechanisms for these tests is discussed. Clinical studies to evaluate the performance of such tests need to follow statistical principles for sound diagnostic test design. Statistical methodology to evaluate such studies can be wide ranging, including receiver operating characteristic (ROC) methodology, logistic regression, discriminant analysis, multiple comparison procedures resampling, Bayesian hierarchical modeling, recursive partitioning, as well as exploratory techniques such as data mining. Recent examples of approved genetic tests are discussed.

Animals↗

Cancer incidence among female flight attendants: a meta-analysis of published data.

BACKGROUND: Flight attendants are exposed to cosmic ionizing radiation and other potential cancer risk factors, but only recently have epidemiological studies been performed to assess the risk of cancer among these workers. The aim of the present work was to evaluate the incidence of various types of cancer among female cabin attendants by combining cancer incidence estimates reported in published studies. METHODS: All follow-up studies reporting standardized incidence ratio (SIR) for cancer among female flight attendants were obtained from online databases and analyzed. A metaanalysis was performed by applying Bayesian hierarchical models, which take into account studies that reported SIR = 0 and natural heterogeneity of study-specific SIRs. RESULTS: A total of seven published studies reporting SIR for several cancer types were extracted. Meta-analysis showed a significant excess of melanoma (meta-SIR 2.15, 95% posterior interval [PI] 1.56-2.88) and breast carcinoma (meta-SIR 1.40; PI 1.19-1.65) and a slight but not significant excess of cancer incidence across types (meta-SIR 1.11, PI 0.98-1.25). CONCLUSIONS: Although further studies are necessary to clarify the exact role of occupational exposure, all airlines should, as some companies do, estimate radiation dose, organize the schedules of crew members in order to reduce further exposure in highly exposed flight attendants, inform crew members about health risks, and give special protection to pregnant women.

Aircraft↗

Rise in malaria incidence rates in South Africa: a small-area spatial analysis of variation in time trends.

Using Bayesian statistical models, the authors investigated spatial and temporal variations in small-area malaria incidence rates for the period mid-1986 to mid-1999 for two districts in northern KwaZulu Natal, South Africa. Maps of spatially smoothed incidence rates at different time points and spatially smoothed time trends in incidence gave a visual impression of the highest increase in incidence occurring where incidence rates previously had been lowest. This was confirmed by conditional autoregressive models, which showed that there was a significant negative association between time trends and smoothed baseline incidence before the steady rise in caseloads began. Growth rates also appeared to be higher in the areas close to the Mozambican border. The main findings of this analysis were that: 1) the spatial distribution of the rise in malaria incidence is uneven and strongly suggests a geographic expansion of high-malaria-risk areas; 2) there is evidence of a stabilization of incidence in areas that had the highest rates before the current escalation of rates began; and 3) areas immediately adjoining the Mozambican border appear to have undergone larger increases in incidence, in contrast to the general pattern of low growth in the more northern, high-baseline-incidence areas, but this was not confirmed by modeling. Smoothing of small-area maps of incidence and growth in incidence (trend) is important for interpretation of the spatial distribution of disease incidence and the spatial distribution of rapid changes in disease incidence.

Bayes Theorem↗

Time trends of breast cancer mortality in Spain during the period 1977-2001 and Bayesian approach for projections during 2002-2016.

BACKGROUND: In recent decades, changes in breast cancer (BC) mortality trends have been observed across Europe. Our objective is to describe BC mortality trend in Spain during 1977-2001 and to estimate BC mortality projection in the period 2002-2016 using a Bayesian approach. MATERIAL AND METHODS: An age-period-cohort (APC) analysis has been carried out in order to investigate the effect of the age, period and birth cohort on BC mortality in Spain during 1977-2001 and to estimate future trends for the period 2002-2016. A Bayesian APC model with an autoregressive structure for the age parameters has been used for projections of BC mortality. RESULTS: BC mortality rates increased 2.18% per year during 1977-1991 followed by a significant fall after 1992 (estimated annual percent change = -2.67%; 95% confidence interval = -2.97, -2.31). Cohorts born before 1952 showed higher risk of death from BC than those born after this year. Projections showed an increase of mortality among women older than 50 years in the period 2002-2016 (range of increase = 10%-40%). CONCLUSIONS: The decrease of BC mortality since 1992 could be attributable to BC down-staging due to early detection and effectiveness of cancer treatment. The effect of ageing on the female population, immigration and the increase of BC incidence observed in Spain could explain the increase in BC mortality predicted for the years to come among women older than 50 years. BC screening to the whole Spanish population and new treatments introduced in the last few years could modify the predictions of BC mortality. Future forecasting studies should be carried out considering these new factors in the natural history of BC in Spain.

Adult↗

Detection of cell-type-specific differentially methylated regions in epigenome-wide association studies.

MOTIVATION: DNA methylation at cytosine-phosphate-guanine (CpG) sites is one of the most important epigenetic markers. Therefore, epidemiologists are interested in investigating DNA methylation in large cohorts through epigenome-wide association studies (EWAS). However, the observed EWAS data are bulk data with signals aggregated from distinct cell types. Deconvolution of cell-type-specific signals from EWAS data is challenging because phenotypes can affect both cell-type proportions and cell-type-specific methylation levels. Recently, there has been active research on detecting cell-type-specific risk CpG sites for EWAS data. However, existing methods all assume that the methylation levels of different CpG sites are independent and perform association detection for each CpG site separately. Although these methods significantly improve the detection at the aggregated-level-identifying a CpG site as a risk CpG site as long as it is associated with the phenotype in any cell type, they have low power in detecting cell-type-specific associations for EWAS with typical sample sizes. RESULTS: Here, we develop a new method, Fine-scale inference for Differentially Methylated Regions (FineDMR), to borrow strengths of nearby CpG sites to improve the cell-type-specific association detection. Via a Bayesian hierarchical model built upon Gaussian process functional regression, FineDMR takes advantage of the spatial dependencies between CpG sites. FineDMR can provide cell-type-specific association detection as well as output subject-specific and cell-type-specific methylation profiles for each subject. Simulation studies and real data analysis show that FineDMR substantially improves the power in detecting cell-type-specific associations for EWAS data. AVAILABILITY AND IMPLEMENTATION: FineDMR is freely available at https://github.com/JiaRuofan/Detection-of-Cell-type-specific-DMRs-in-EWAS.

DNA Methylation↗

Bayesian estimation of allele-specific expression in the presence of phasing uncertainty.

MOTIVATION: Allele-specific expression (ASE) analyses aim to detect imbalanced expression of maternal versus paternal copies of an autosomal gene. Such allelic imbalance can result from a variety of cis-acting causes, including disruptive mutations within one copy of a gene that impact the stability of transcripts, as well as regulatory variants outside the gene that impact transcription initiation. Current methods for ASE estimation suffer from a number of shortcomings, such as relying on only one variant within a gene, assuming perfect phasing information across multiple variants within a gene, or failing to account for alignment biases and possible genotyping errors. RESULTS: We developed BEASTIE, a Bayesian hierarchical model designed for precise ASE quantification at the gene level, based on given genotypes and RNA-Seq data. BEASTIE addresses the complexities of allelic mapping bias, genotyping error, and phasing errors by incorporating empirical phasing error rates derived from Genome-in-a-Bottle individual NA12878. BEASTIE surpasses existing methods in accuracy, especially in scenarios with high phasing errors. This improvement is critical for identifying rare genetic variants often obscured by such errors. Through rigorous validation on simulated data and application to real data from the 1000 Genomes Project, we establish the robustness of BEASTIE. These findings underscore the value of BEASTIE in revealing patterns of ASE across gene sets and pathways. AVAILABILITY AND IMPLEMENTATION: The software is freely available from Github (https://github.com/x811zou/BEASTIE); and Zendo (DOI: 10.5281/zenodo.15062124).

Bayes Theorem↗

Bayesian reconstruction and differential testing of excised introns.

MOTIVATION: Characterizing the differential excision of introns is critical for understanding the functional complexity of a cell or tissue, from normal developmental processes to disease pathogenesis. Most transcript reconstruction methods infer full-length transcripts from high-throughput sequencing data. However, this is a challenging task due to incomplete annotations and the heterogeneous expression of transcripts across cell-types, tissues, and experimental conditions. Several recent methods circumvent these difficulties by considering local splicing events, but these methods lose transcript-level splicing information and may conflate similar, but distinct transcripts. RESULTS: In this work, we formalize a new transcript reconstruction problem that interpolates between the full-length and local splicing perspectives by considering sequences of exon-exon junctions (SEEJs) that co-occur in transcripts. We then present a hierarchical Bayesian admixture model and posterior inference algorithms for computing SEEJs (BSEEJ), and a generalized linear model for characterizing differential SEEJ usage based on model parameter estimates. We show that BSEEJ achieves high F1 score for reconstruction tasks and improved accuracy and sensitivity in differential splicing when compared with six transcript and local splicing methods on simulated data. Lastly, we evaluate BSEEJ on experimental data based on transcript reconstruction, novelty of transcripts produced, model sensitivity to hyperparameters, and a functional analysis of differentially expressed SEEJs. AVAILABILITY AND IMPLEMENTATION: BSEEJ is freely available at https://github.com/bayesomicslab/BSEEJ.

Bayes Theorem↗

BAPS 2: enhanced possibilities for the analysis of genetic population structure.

UNLABELLED: Bayesian statistical methods based on simulation techniques have recently been shown to provide powerful tools for the analysis of genetic population structure. We have previously developed a Markov chain Monte Carlo (MCMC) algorithm for characterizing genetically divergent groups based on molecular markers and geographical sampling design of the dataset. However, for large-scale datasets such algorithms may get stuck to local maxima in the parameter space. Therefore, we have modified our earlier algorithm to support multiple parallel MCMC chains, with enhanced features that enable considerably faster and more reliable estimation compared to the earlier version of the algorithm. We consider also a hierarchical tree representation, from which a Bayesian model-averaged structure estimate can be extracted. The algorithm is implemented in a computer program that features a user-friendly interface and built-in graphics. The enhanced features are illustrated by analyses of simulated data and an extensive human molecular dataset. AVAILABILITY: Freely available at http://www.rni.helsinki.fi/~jic/bapspage.html.

Algorithms↗

Bayesian search of functionally divergent protein subgroups and their function specific residues.

MOTIVATION: The rapid increase in the amount of protein sequence data has created a need for an automated identification of evolutionarily related subgroups from large datasets. The existing methods typically require a priori specification of the number of putative groups, which defines the resolution of the classification solution. RESULTS: We introduce a Bayesian model-based approach to simultaneous identification of evolutionary groups and conserved parts of the protein sequences. The model-based approach provides an intuitive and efficient way of determining the number of groups from the sequence data, in contrast to the ad hoc methods often exploited for similar purposes. Our model recognizes the areas in the sequences that are relevant for the clustering and regards other areas as noise. We have implemented the method using a fast stochastic optimization algorithm which yields a clustering associated with the estimated maximum posterior probability. The method has been shown to have high specificity and sensitivity in simulated and real clustering tasks. With real datasets the method also highlights the residues close to the active site. AVAILABILITY: Software 'kPax' is available at http://www.rni.helsinki.fi/jic/softa.html

Algorithms↗

Combining longitudinal studies of PSA.

Prostate-specific antigen (PSA) is a biomarker commonly used to screen for prostate cancer. Several studies have examined PSA growth rates prior to prostate cancer diagnosis. However, the resulting estimates are highly variable. In this article we propose a non-linear Bayesian hierarchical model to combine longitudinal data on PSA growth from three different studies. Our model enables novel investigations into patterns of PSA growth that were previously impossible due to sample size limitations. The goals of our analysis are twofold: first, to characterize growth rates of PSA accounting for differences when combining data from different studies; second, to investigate the impact of clinical covariates such as advanced disease and unfavorable histology on PSA growth rates.

Adult↗

Bayesian analysis of mutational spectra.

Studies that examine both the frequency of gene mutation and the pattern or spectrum of mutational changes can be used to identify chemical mutagens and to explore the molecular mechanisms of mutagenesis. In this article, we propose a Bayesian hierarchical modeling approach for the analysis of mutational spectra. We assume that the total number of independent mutations and the numbers of mutations falling into different response categories, defined by location within a gene and/or type of alteration, follow binomial and multinomial sampling distributions, respectively. We use prior distributions to summarize past information about the overall mutation frequency and the probabilities corresponding to the different mutational categories. These priors can be chosen on the basis of data from previous studies using an approach that accounts for heterogeneity among studies. Inferences about the overall mutation frequency, the proportions of mutations in each response category, and the category-specific mutation frequencies can be based on posterior distributions, which incorporate past and current data on the mutant frequency and on DNA sequence alterations. Methods are described for comparing groups and for assessing dose-related trends. We illustrate our approach using data from the literature.

Animals↗

Genome-wide estimation of transcript concentrations from spotted cDNA microarray data.

A method providing absolute transcript concentrations from spotted microarray intensity data is presented. Number of transcripts per microg total RNA, mRNA or per cell, are obtained for each gene, enabling comparisons of transcript levels within and between tissues. The method is based on Bayesian statistical modelling incorporating available information about the experiment from target preparation to image analysis, leading to realistically large confidence intervals for estimated concentrations. The method was validated in experiments using transcripts at known concentrations, showing accuracy and reproducibility of estimated concentrations, which were also in excellent agreement with results from quantitative real-time PCR. We determined the concentration for 10,157 genes in cervix cancers and a pool of cancer cell lines and found values in the range of 10(5)-10(10) transcripts per microg total RNA. The precision of our estimates was sufficiently high to detect significant concentration differences between two tumours and between different genes within the same tumour, comparisons that are not possible with standard intensity ratios. Our method can be used to explore the regulation of pathways and to develop individualized therapies, based on absolute transcript concentrations. It can be applied broadly, facilitating the construction of the transcriptome, continuously updating it by integrating future data.

Bayes Theorem↗

Protease mutation M89I/V is linked to therapy failure in patients infected with the HIV-1 non-B subtypes C, F or G.

OBJECTIVE: To investigate whether and how mutations at position 89 of HIV-1 protease were associated with protease inhibitor (PI) failure, and what is the impact of the HIV-1 subtype. METHODS: In a database containing pol nucleotide sequences and treatment history, the correlation between PI experience and mutations at codon 89 was determined separately for subtype B and several non-B subtypes. A Bayesian network model was used to map the resistance pathways in which M89I/V is involved for subtype G. The phenotypic effect of M89I/V for several PIs was also measured. RESULTS: The analysis showed that for the subtypes C, F and G in which the wild-type codon at 89 was M compared to L for subtype B, M89I/V was significantly more frequently observed in PI-treated patients displaying major resistance mutations to PIs than in drug-naive patients. M89I/V was strongly associated with PI resistance mutations at codons 71, 74 and 90. Phenotypically, M89I/V alone did not confer a reduced susceptibility to PIs. However, when combined with L90M, a significantly reduced susceptibility to nelfinavir was observed (P < 0.05) in comparison with strains with L90M alone. CONCLUSIONS: The results of the present study show that M89I/V is associated with PI experience in subtypes C, F and G but not in subtype B. M89I/V should be considered a secondary PI mutation with an important effect on nelfinavir susceptibility in the presence of L90M.

Adult↗

Profiling providers on use of adjuvant chemotherapy by combining cancer registry and medical record data.

PURPOSE: Treatment information collected by cancer registries can be used to monitor the provision of guideline-recommended chemotherapy to colorectal cancer patients. Incomplete information may bias comparisons of these rates. We developed statistical methods that combine data from a registry and physicians' records to assess hospital quality. DATA: From California Cancer Registry data, we selected all patients (n=12,594) newly diagnosed with stage III colon cancer or stage II or III rectal cancer from 428 hospitals during the years 1994 to 1998. To assess rates and predictors of underreporting of chemotherapy, we surveyed physicians treating 1449 of these patients from 98 hospitals during the years 1996 to 1997. METHODS: Using Bayesian statistical models, we imputed unobserved treatments. We studied the impact of underreporting on provider profiling by comparing rankings, estimates, and credible intervals based only on registry data to those incorporating physician survey data. RESULTS: Analyses that account for incompleteness of reporting yielded wider credible intervals for provider profiles than those that ignored such incompleteness. Among the 109 (25%) hospitals in the highest quartile of chemotherapy rates according to the registry data, 16 were not so classified when incomplete reporting was taken into account. With the more comprehensive model, 12 hospitals could be identified that ranked in the top quartile with probability>0.90. CONCLUSION: Estimates of adjusted hospital chemotherapy rates based solely on cancer registry data overstate the precision of assessments of hospital quality. Using additional information from a physician survey and applying rigorous statistical models, better inferences can be drawn about provider quality.

Adolescent↗

Neonatal intensive care unit characteristics affect the incidence of severe intraventricular hemorrhage.

OBJECTIVES: The incidence of intraventricular hemorrhage (IVH), adjusted for known risk factors, varies across neonatal intensive care units (NICU)s. The effect of NICU characteristics on this variation is unknown. The objective was to assess IVH attributable risks at both patient and NICU levels. STUDY DESIGN: Subjects were <33 weeks' gestation, <4 days old on admission in the Canadian Neonatal Network database (all infants admitted in 1996-97 to 17 NICUs). The variation in severe IVH rates was analyzed using Bayesian hierarchical modeling for patient level and NICU level factors. RESULTS: Of 3772 eligible subjects, the overall crude incidence rates of grade 3-4 IVH was 8.3% (NICU range 2.0-20.5%). Male gender, extreme preterm birth, low Apgar score, vaginal birth, outborn birth, and high admission severity of illness accounted for 30% of the severe IVH rate variation; admission day therapy-related variables (treatment of acidosis and hypotension) accounted for an additional 14%. NICU characteristics, independent of patient level risk factors, accounted for 31% of the variation. NICUs with high patient volume and high neonatologist/staff ratio had lower rates of severe IVH. CONCLUSIONS: The incidence of severe IVH is affected by NICU characteristics, suggesting important new strategies to reduce this important adverse outcome.

Acute Disease↗

Identifying the structure in cuttlefish visual signals.

The common cuttlefish (Sepia officinalis) communicates and camouflages itself by changing its skin colour and texture. Hanlon and Messenger (1988 Phil. Trans. R. Soc. Lond. B 320, 437-487) classified these visual displays, recognizing 13 distinct body patterns. Although this conclusion is based on extensive observations, a quantitative method for analysing complex patterning has obvious advantages. We formally define a body pattern in terms of the probabilities that various skin features are expressed, and use Bayesian statistical methods to estimate the number of distinct body patterns and their visual characteristics. For the dataset of cuttlefish coloration patterns recorded in our laboratory, this statistical method identifies 12-14 different patterns, a number consistent with the 13 found by Hanlon and Messenger. If used for signalling these would give a channel capacity of 3.4 bits per pattern. Bayesian generative models might be useful for objectively describing the structure in other complex biological signalling systems.

Animal Communication↗

A bimodal pattern of relatedness between the Salmonella Paratyphi A and Typhi genomes: convergence or divergence by homologous recombination?

All Salmonella can cause disease but severe systemic infections are primarily caused by a few lineages. Paratyphi A and Typhi are the deadliest human restricted serovars, responsible for approximately 600,000 deaths per annum. We developed a Bayesian changepoint model that uses variation in the degree of nucleotide divergence along two genomes to detect homologous recombination between these strains, and with other lineages of Salmonella enterica. Paratyphi A and Typhi showed an atypical and surprising pattern. For three quarters of their genomes, they appear to be distantly related members of the species S. enterica, both in their gene content and nucleotide divergence. However, the remaining quarter is much more similar in both aspects, with average nucleotide divergence of 0.18% instead of 1.2%. We describe two different scenarios that could have led to this pattern, convergence and divergence, and conclude that the former is more likely based on a variety of criteria. The convergence scenario implies that, although Paratyphi A and Typhi were not especially close relatives within S. enterica, they have gone through a burst of recombination involving more than 100 recombination events. Several of the recombination events transferred novel genes in addition to homologous sequences, resulting in similar gene content in the two lineages. We propose that recombination between Typhi and Paratyphi A has allowed the exchange of gene variants that are important for their adaptation to their common ecological niche, the human host.

Algorithms↗