Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “statistical inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

Surgical utilization statistics: some methodologic considerations.

This article considers variations in the recording and counting protocols used in the generation of surgical utilization data. A single raw data source was manipulated to reproduce several common protocols to illustrate the statistically significant differences that can result in the volume of surgical utilization considered for both individual procedures and groups of procedures. The results suggest that if recording or counting protocols differ in the samples under consideration, comparison of the statistics and inferences drawn as to utilization therein may be confounded. It is quite possible, therefore, that some of the results of earlier surgical utilization studies may be confounded by such differences in protocols. While there may be valid differences in surgical utilization across different settings, our findings suggest that until methods used in previous work are investigated and reconciled, caution should be exercised in the utilization of this research in public policy making.

Female↗

Mixed graphical models for simultaneous model identification and control applied to the glucose-insulin metabolism.

In this paper a method for model identification of biological systems described by stochastic linear differential equations using a new computational technique for statistical Bayesian inference, namely mixed graphical models in the sense of Lauritzen and Wermuth, is presented. The model is identified in terms of biological model parameters and noise parameters. This non-linear estimation problem is solved by means of an exact inference algorithm. The parameter estimates are given as a-posteriori distributions which can be interpreted as fuzzy possibility distributions. For model-based simulations of the underlying biological system the model parameters are represented as uncertain parameters with the distributions obtained from the estimation procedure. We apply the presented methods to a model for the glucose-insulin metabolism: the Karlsburg model for type I diabetes.

Bayes Theorem↗

Flexible sequence similarity searching with the FASTA3 program package.

The FASTA3 and FASTA2 packages provide a flexible set of sequence-comparison programs that are particularly valuable because of their accurate statistical estimates and high-quality alignments. Traditionally, sequence similarity searches have sought to ask one question: "Is my query sequence homologous to anything in the database?" Both FASTA and BLAST can provide reliable answers to this question with their statistical estimates; if the expectation value E is < 0.001-0.01 and you are not doing hundreds of searches a day, the answer is probably yes. In general, the most effective search strategies follow these rules: 1. Whenever possible, compare at the amino acid level, rather than the nucleotide level. Search first with protein sequences (blastp, fasta3, and ssearch3), then with translated DNA sequences (fastx, blastx), and only at the DNA level as a last resort (Table 5). 2. Search the smallest database that is likely to contain the sequence of interest (but it must contain many unrelated sequences for accurate statistical estimates). 3. Use sequence statistics, rather than percent identity or percent similarity, as your primary criterion for sequence homology. 4. Check that the statistics are likely to be accurate by looking for the highest-scoring unrelated sequence, using prss3 to confirm the expectation, and searching with shuffled copies of the query sequence [randseq, searches with shuffled sequences should have E approx 1.0]. 5. Consider searches with different gap penalties and other scoring matrices. Searches with long query sequences against full-length sequence libraries will not change dramatically when BLOSUM62 is used instead of BLOSUM50 (20), or a gap penalty of -14/-2 is used in place of -12/-2. However, shallower or more stringent scoring matrices are more effective at uncovering relationships in partial sequences (3,18), and they can be used to sharpen dramatically the scope of the similarity search. However, as illustrated in the last section, the E value is only the first step in characterizing a sequence relationship. Once one has confidence that the sequences are homologous, one should look at the sequence alignments and percent identities, particularly when searching with lower quality sequences. When sequence alignments are very short, the alignment should become more significant when a shallower scoring matrix is used, e.g., BLOSUM62 rather than BLOSUM50 (remember to change the gap penalties). Homology can be reliably inferred from statistically significant similarity. Whereas homology implies common three-dimensional structure, homology need not imply common function. Orthologous sequences usually have similar functions, but paralogous sequences often acquire very different functional roles. Motif databases, such as PROSITE (21), can provide evidence for the conservation of critical functional residues. However, motif identity in the absence of overall sequence similarity is not a reliable indicator of homology.

Amino Acid Sequence↗

Sampling design, response rates, and analysis weights for the National Human Exposure Assessment Survey (NHEXAS) in EPA region 5.

For the Phase I field test of the National Human Exposure Assessment Survey (NHEXAS) in U.S. Environmental Protection Agency (EPA) Region 5, this paper presents the survey sampling design, the response rates achieved, and the sample weighting procedure implemented to compensate for unit nonresponse. To enable statistically defensible inferences to the entire region, a sample of about 250 members of the household population in EPA Region 5 was selected using a stratified multistage probability-based survey sampling design. Sample selection proceeded in four nested stages: (1) sample counties; (2) area segments based on Census blocks within sample counties; (3) housing units (HUs) within sample segments; and (4) individual participants within sample households. Each fourth-stage sample member was asked to participate in 6 days of exposure monitoring. A subsample of participants was asked to participate in two rounds of longitudinal follow-up data collection. Approximately 70% of all sample households participated in household screening interviews in which rosters of household members were developed. Over 70% of the sample subjects selected from these households completed the Baseline Questionnaire regarding their demographic characteristics and potential for exposures. And, over 75% of these sample members went on to complete at least the core environmental monitoring, including personal exposures to volatile organic compounds (VOCs) and tap water concentrations of metals. The sample weighting procedures used the data collected in the screening interviews for all household members to fit logistic models for nonresponse in the later phases of the study. Moreover, the statistical analysis weights were poststratified to 1994 State population projections obtained from the Bureau of the Census to ensure consistency with other statistics for the Region.

Adolescent↗

Reconstructing the early spatial spread of pandemic respiratory viruses in the United States.

Understanding the geographic spread of emerging respiratory viruses is critical for pandemic preparedness, yet the early spatiotemporal dynamics of the 2009 H1N1 pandemic influenza and severe acute respiratory syndrome coronavirus 2 in the United States remain unclear. While mobility and genomic data have revealed important aspects of pandemic spatial spread, several key questions remain: Did the two pandemics follow similar spatial transmission routes? How rapidly did they spread across the United States? What role did stochastic processes play in early spatial transmission? To address these questions, we integrated high-resolution disease data with a robust, data-efficient inference framework combining air travel, commuting flows, and pathogen superspreading potentials to reconstruct their spatial spread across US metropolitan areas. The two pandemics exhibited distinct transmission pathways across locations; however, both pandemics established local circulation in most metropolitan areas within weeks, driven by several shared transmission hubs. Early spatial spread was more strongly associated with air travel than with commuting, though stochastic dynamics introduced substantial uncertainty in transmission routes, creating challenges for timely detection and control. Simulations indicate that broad wastewater surveillance coverage beyond top transmission hubs coupled with effective infection control may slow initial spatial expansion. Our findings highlight the rapid, stochastic spread of pandemic respiratory pathogens and the difficulties of early outbreak containment.

Humans↗

Statistical methods of estimation and inference for functional MR image analysis.

Two questions arising in the analysis of functional magnetic resonance imaging (fMRI) data acquired during periodic sensory stimulation are: i) how to measure the experimentally determined effect in fMRI time series; and ii) how to decide whether an apparent effect is significant. Our approach is first to fit a time series regression model, including sine and cosine terms at the (fundamental) frequency of experimental stimulation, by pseudogeneralized least squares (PGLS) at each pixel of an image. Sinusoidal modeling takes account of locally variable hemodynamic delay and dispersion, and PGLS fitting corrects for residual or endogenous autocorrelation in fMRI time series, to yield best unbiased estimates of the amplitudes of the sine and cosine terms at fundamental frequency; from these parameters the authors derive estimates of experimentally determined power and its standard error. Randomization testing is then used to create inferential brain activation maps (BAMs) of pixels significantly activated by the experimental stimulus. The methods are illustrated by application to data acquired from normal human subjects during periodic visual and auditory stimulation.

Acoustic Stimulation↗

Geometric cooperativity and anticooperativity of three-body interactions in native proteins.

Characterizing multibody interactions of hydrophobic, polar, and ionizable residues in protein is important for understanding the stability of protein structures. We introduce a geometric model for quantifying 3-body interactions in native proteins. With this model, empirical propensity values for many types of 3-body interactions can be reliably estimated from a database of native protein structures, despite the overwhelming presence of pairwise contacts. In addition, we define a nonadditive coefficient that characterizes cooperativity and anticooperativity of residue interactions in native proteins by measuring the deviation of 3-body interactions from 3 independent pairwise interactions. It compares the 3-body propensity value from what would be expected if only pairwise interactions were considered, and highlights the distinction of propensity and cooperativity of 3-body interaction. Based on the geometric model, and what can be inferred from statistical analysis of such a model, we find that hydrophobic interactions and hydrogen-bonding interactions make nonadditive contributions to protein stability, but the nonadditive nature depends on whether such interactions are located in the protein interior or on the protein surface. When located in the interior, many hydrophobic interactions such as those involving alkyl residues are anticooperative. Salt-bridge and regular hydrogen-bonding interactions, such as those involving ionizable residues and polar residues, are cooperative. When located on the protein surface, these salt-bridge and regular hydrogen-bonding interactions are anticooperative, and hydrophobic interactions involving alkyl residues become cooperative. We show with examples that incorporating 3-body interactions improves discrimination of protein native structures against decoy conformations. In addition, analysis of cooperative 3-body interaction may reveal spatial motifs that can suggest specific protein functions.

Computer Simulation↗

A goodness-of-fit approach to inference procedures for the kappa statistic: confidence interval construction, significance-testing and sample size estimation.

We propose a new procedure for constructing a confidence interval about the kappa statistic in the case of two raters and a dichotomous outcome. The procedure is based on a chi-square goodness-of-fit test as applied to a model frequently used for clustered binary data. The procedure provides coverage levels that are accurate in samples of smaller size than those required for other procedures. The procedure also has use for significance-testing and the planning of corresponding sample size requirements.

Confidence Intervals↗

Contaminants in fish of the Hackensack Meadowlands, New Jersey: size, sex, and seasonal relationships as related to health risks.

The trace metal content and related safety (health risk) of Hackensack River fish were assessed within the Hackensack Meadowlands of New Jersey, USA. Eight elements were analyzed in the edible portion (i.e., muscle) of species commonly taken by anglers in the area. The white perch collection (Morone americana) was large enough (n = 168) to enable statistically significant inferences, but there were too few brown bullheads and carp to reach definite conclusions. Of the eight elements analyzed, the one that accumulates to the point of being a health risk in white perch is mercury (Hg). Relationships between mercury concentrations and size and with collection season were observed; correlation with lipid content, total polychlorinated biphenyl (PCB) content, or collection site were very weak. Only 18% of the Hg was methylated in October (n = 8), whereas June and July fish (n = 12) had 100% methylation of Hg. White perch should not be considered edible because the Hg level exceeded the "one meal per month" action level of 0.47 microg/g wet weight (ppm) in 32% of our catch and 2.5% exceeded the "no consumption at all" level of 1 microg/g. The larger fish represent greater risk for Hg. Furthermore, the warmer months, when more recreational fishing takes place, might present greater risk. A more significant reason for avoiding white perch is the PCB contamination because 40% of these fish exceeded the US Food and Drug Administration (FDA) action level of 2000 ng/g for PCBs and all white perch exceeded the US Environmental Protection Agency cancer/health guideline (49 ng/g) of no more than one meal/month. In fact, nearly all were 10 times that advisory level. There were differences between male and female white perch PCB levels, with nearly all of those above the US FDA action level being male. Forage fish (mummichogs and Atlantic silversides) were similarly analyzed, but no correlations were found with any other parameters. The relationship of collection site to contaminants cannot be demonstrated because sufficient numbers of game fish could not be collected at many sites at all seasons.

Animals↗

Security considerations for present and future medical databases.

In this paper we consider the security of medical databases. We give an overview of the security problems and the possible available mechanisms for the prevention of security compromises. Many of the security problems are common to all databases. However, the problem of data inference from statistical queries is particularly pertinent to medical databases and consequently we treat this problem in more detail. The paper concludes with a proposal for a Security Subsystem in a database management system.

Computer Security↗

Public and private hypotheses.

The uses of hypothesis in scientific medicine and in clinical medicine are broadly similar; but they differ in subtle but important logical details. The key distinction is that scientific medicine deals with broad ("public") hypotheses about entire populations; and the predominant problem of inference is statistics in the face of epistemological uncertainty. In contrast, clinical medicine deals with individual, tailor-made ("private") hypotheses; the uncertainty (e.g. the diagnosis or recurrence risk in relatives) is ontological, and the main feature of the analysis is probability. The same logical rules are at times mistakenly applied to both. Disregard of these subtleties may lead to paradoxes, contradictions, and at times diastrously false conclusion. The basic principles and distinctions are laid out with illustrations that are most readily provided from medical genetics, but apply to all branches of medicine.

Epidemiologic Methods↗

Engineering of large numbers of highly specific homing endonucleases that induce recombination on novel DNA targets.

The last decade has seen the emergence of a universal method for precise and efficient genome engineering. This method relies on the use of sequence-specific endonucleases such as homing endonucleases. The structures of several of these proteins are known, allowing for site-directed mutagenesis of residues essential for DNA binding. Here, we show that a semi-rational approach can be used to derive hundreds of novel proteins from I-CreI, a homing endonuclease from the LAGLIDADG family. These novel endonucleases display a wide range of cleavage patterns in yeast and mammalian cells that in most cases are highly specific and distinct from I-CreI. Second, rules for protein/DNA interaction can be inferred from statistical analysis. Third, novel endonucleases can be combined to create heterodimeric protein species, thereby greatly enhancing the number of potential targets. These results describe a straightforward approach for engineering novel endonucleases with tailored specificities, while preserving the activity and specificity of natural homing endonucleases, and thereby deliver new tools for genome engineering.

Amino Acid Sequence↗

Molecular epidemiology and phylogeographic architecture of oncogenic intracellular bacteria in cervical cancer patients across Northern China.

BACKGROUND: Oncogenic intracellular bacteria, including Chlamydia trachomatis, Mycoplasma genitalium, and Fusobacterium nucleatum, have emerged as significant contributors to cervical carcinogenesis. Despite growing interest in microbial oncology, the molecular epidemiological landscape and phylogeographic distribution of these pathogens in Northern China remain poorly characterized. This study aimed to determine the prevalence, co-infection patterns, genotypic diversity, and spatial phylogeographic clustering of oncogenic intracellular bacteria among cervical cancer patients across five provinces of Northern China. METHODS: A cross-sectional, multi-center study was conducted between March 2022 and November 2024 across Shaanxi, Heilongjiang, Beijing, Shandong, and Inner Mongolia. Cervical swab specimens were collected from 1247 confirmed cervical cancer patients. Pathogen detection was performed using multiplex real-time polymerase chain reaction, 16S rRNA gene amplicon sequencing, and whole-genome sequencing. Phylogeographic analyses employed maximum likelihood and Bayesian evolutionary inference frameworks. Statistical analyses included multivariate logistic regression and geographic information system-based spatial clustering. RESULTS: The overall prevalence of at least one oncogenic intracellular bacterium was 68.3% (n&#xa0;=&#xa0;852). Chlamydia trachomatis was the most prevalent pathogen detected in 41.2% of participants. Co-infection with two or more bacteria was identified in 29.7% of cases and was independently associated with advanced-stage cervical cancer (adjusted odds ratio&#xa0;=&#xa0;2.87; 95% confidence interval: 1.94 to 4.23; p&#xa0;<&#xa0;0.001). Phylogeographic analysis revealed three distinct molecular clades with evidence of bidirectional gene flow between Shaanxi and Heilongjiang. Whole-genome sequencing identified 14 novel virulence gene variants not previously characterized in Chinese clinical isolates. CONCLUSIONS: Oncogenic intracellular bacteria are highly prevalent and genotypically diverse among cervical cancer patients in Northern China. The identified phylogeographic clustering and novel virulence variants have direct implications for regional screening programs, targeted antimicrobial strategies, and the development of region-specific molecular diagnostic panels.

Cervical cancer↗

Interdependence and distribution of subclinical mastitis and intramammary infection among udder quarters in dairy cattle.

The objective of the present study was to determine the distribution of subclincial mastitis and intramammary infection (IMI) across udder quarters and to quantify, if any, the level of interdependence between quarters. Subclinical mastitis was defined as an udder quarter somatic cell count of > or =250,000. Intramammary infection was defined as the presence of pathogens in an udder quarter. The data used in the analyses consisted 6577 test-day records from 2034 lactations with observations on subclinical mastitis and 8428 test-day records from 2575 lactations with observations on IMI; each test-day record consisted data on all four udder quarters. Records were from animals on three research herds in southern Ireland. Observed distributions of IMI were compared to expected distributions, estimated using a binomial probability distribution, assuming independence among quarters. Generally, the observed frequencies deviated significantly from the expected frequencies, indicating some level of interdependence between quarters. Intraclass correlations were estimated from a mixed model including cow and test-day record nested within cow as random effects; fixed effects in the model included herd, year of calving, month of calving, parity, days in milk at sampling and interactions. Intraclass correlations within test-day record ranged from 0.14 to 0.50 for subclinical mastitis and from 0.02 to 0.40 for IMI. Thus, substantial interdependence between quarters exists which should be accounted for when analysing data from experiments across udder quarters. Failure to do so may result in incorrect inferences from statistical tests.

Animals↗

Further use of nearly complete 28S and 18S rRNA genes to classify Ecdysozoa: 37 more arthropods and a kinorhynch.

This work expands on a study from 2004 by Mallatt, Garey, and Shultz [Mallatt, J.M., Garey, J.R., Shultz, J.W., 2004. Ecdysozoan phylogeny and Bayesian inference: first use of nearly complete 28S and 18S rRNA gene sequences to classify the arthropods and their kin. Mol. Phylogenet. Evol. 31, 178-191] that evaluated the phylogenetic relationships in Ecdysozoa (molting animals), especially arthropods. Here, the number of rRNA gene-sequences was effectively doubled for each major group of arthropods, and sequences from the phylum Kinorhyncha (mud dragons) were also included, bringing the number of ecdysozoan taxa to over 80. The methods emphasized maximum likelihood, Bayesian inference and statistical testing with parametric bootstrapping, but also included parsimony and minimum evolution. Prominent findings from our combined analysis of both genes are as follows. The fundamental subdivisions of Hexapoda (insects and relatives) are Insecta and Entognatha, with the latter consisting of collembolans (springtails) and a clade of proturans plus diplurans. Our rRNA-gene data provide the strongest evidence to date that the sister group of Hexapoda is Branchiopoda (fairy shrimps, tadpole shrimps, etc.), not Malacostraca. The large, Pancrustacea clade (hexapods within a paraphyletic Crustacea) divided into a few basic subclades: hexapods plus branchiopods; cirripedes (barnacles) plus malacostracans (lobsters, crabs, true shrimps, isopods, etc.); and the basally located clades of (a) ostracods (seed shrimps) and (b) branchiurans (fish lice) plus the bizarre pentastomids (tongue worms). These findings about Pancrustacea agree with a recent study by Regier, Shultz, and Kambic that used entirely different genes [Regier, J.C., Shultz, J.W., Kambic, R.E., 2005a. Pancrustacean phylogeny: hexapods are terrestrial crustaceans and maxillopods are not monophyletic. Proc. R. Soc. B 272, 395-401]. In Malacostraca, the stomatopod (mantis shrimp) was not at the base of the eumalacostracans, as is widely claimed, but grouped instead with an euphausiacean (krill). Within centipedes, Craterostigmus was the sister to all other pleurostigmophorans, contrary to the consensus view. Our trees also united myriapods (millipedes and centipedes) with chelicerates (horseshoe crabs, spiders, scorpions, and relatives) and united pycnogonids (sea spiders) with chelicerates, but with much less support than in the previous rRNA-gene study. Finally, kinorhynchs joined priapulans (penis worms) at the base of Ecdysozoa.

Animals↗

Emergency contraception: randomized comparison of advance provision and information only.

OBJECTIVE: To determine whether multiple courses of emergency contraceptive therapy supplied in advance of need would tempt women using barrier methods to take risks with their more effective ongoing contraceptive methods. METHODS: We randomly assigned 411 condom users attending an urban family planning clinic in Pune, India, to receive either information about emergency contraception along with three courses of therapy to keep in case of need, or to receive only information, including that about the locations where they could obtain emergency contraception if needed. For up to 1 year, women returned quarterly for follow-up, answering questions about unprotected intercourse, emergency contraceptive use, pregnancies, sexually transmitted infections, and acceptability. RESULTS: Women given advance supplies reported unprotected intercourse at rates nearly identical to those among women given only information (0.012 versus 0.016 acts per month). Among those who did have unprotected intercourse, however, supply recipients were nearly twice as likely (79% versus 44%) to have taken emergency contraception, although numbers were too small to permit statistically significant inferences. No women used emergency contraception more than once during the study, even though everyone in the advance-supplies group had extra doses available. All women found knowing about emergency contraception useful, and all those receiving only information wished they had received supplies as well. CONCLUSION: Multiple emergency contraception doses supplied in advance did not tempt condom users to risk unprotected intercourse. After unprotected intercourse, however, those with pills on hand used them more often. Women found advance provision useful.

Adult↗

Downstream outcomes: using insurance claims data to screen for errors in clinical laboratory testing.

A methodology is described by which health insurance claims data might be used to discover the occurrence of systematic errors by clinical laboratories. False-positive results should generate a series of tests or treatments that are eventually abandoned as the false signal of the initial test is discovered while false-negative results may cause necessary tests or treatments to be unduly delayed. False results may also generate adverse outcomes such as an unusually high number of deaths or hospitalizations among persons who have received particular laboratory tests. Health insurance claims data may be used to discover these patterns and how the inclusion of laboratory results on claims would improve the precision of such inferences. Appropriate statistical tests are discussed.

Centers for Medicare and Medicaid Services, U.S.↗