Search PubMedSearch

SEARCH · Search PubMed

Results for “Uncertainty score”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

scATAnno: Automated Cell Type Annotation for Single-cell ATAC-seq Data.

Recent advances in single-cell epigenomic techniques have increased the demand for single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) analysis. One key analytical task is to determine cell type identity based on epigenetic data. Here, we introduce scATAnno, a Python package designed to automatically annotate scATAC-seq data using large-scale scATAC-seq reference atlases. This workflow generates reference atlases from publicly available datasets, enabling accurate cell type annotation by integrating query data with reference atlases without the use of single-cell RNA sequencing (scRNA-seq) data. To enhance annotation accuracy, we incorporated k-nearest neighbors (KNN)-based and weighted distance-based uncertainty scores to effectively detect cell populations within the query data that are distinct from all cell types in the reference data. We compared and benchmarked scATAnno against five other published cell annotation approaches, demonstrating its superior performance across multiple datasets and metrics. We further showcased the utility of scATAnno across multiple datasets, including peripheral blood mononuclear cells (PBMCs), triple-negative breast cancer (TNBC), and basal cell carcinoma (BCC), and demonstrated that scATAnno accurately annotates cell types across diverse biological conditions. Overall, scATAnno is a useful tool for scATAC-seq reference atlas construction and cell type annotation and can facilitate the interpretation of new scATAC-seq datasets in complex biological systems. scATAnno is publicly available at https://scatanno-main.readthedocs.io/.

Single-Cell Analysis

Stools containing altered blood-plasma urea: creatinine ratio as a simple test for the source of bleeding.

The plasma urea:creatinine ratio (U:C ratio) is known to be elevated in cases of upper gastrointestinal bleeding. Almost all patients with haematemesis have upper gastrointestinal (or generalized) bleeding so that in this study we characterized the diagnostic power of the U:C ratio in patients with stools containing altered blood without haematemesis in the hope that this simple laboratory test (used in conjunction, perhaps, with clinical data) might reduce the number of patients subjected to an unrewarding gastroscopy or colonoscopy. Of 76 cases seen in a provincial and a metropolitan hospital, 42 and 34 patients had upper and lower gastrointestinal bleeding, respectively. Fifty-four per cent of those with upper gastrointestinal bleeding and none of those with lower gastrointestinal bleeding had U:C ratios above 110 on admission. However, a discriminating level of 90 is considered to be more suitable, judged by the quadratic uncertainty score. At this level the odds for upper gastrointestinal bleeding were 15:1.

Aged

raxtax: a k-mer-based non-Bayesian taxonomic classifier.

MOTIVATION: Taxonomic classification in biodiversity studies is the process of assigning the anonymous sequences of a marker gene (barcode) or whole genomes (metagenomics) to a specific lineage using a reference database that contains named sequences in a known taxonomy. This classification is important for assessing the diversity of biological systems. Taxonomic classification faces two main challenges: first, accuracy is critical as errors can propagate to downstream analysis results; and second, the classification time requirements can limit study size and study design, in particular when considering the constantly growing reference databases. To address these two challenges, we introduce raxtax, an efficient, novel taxonomic classification tool for barcodes that uses common k-mers between all pairs of query and reference sequences. We also introduce two novel uncertainty scores which take into account the fundamental biases of reference databases. RESULTS: We validate raxtax on three widely-used empirical reference databases and show that it is 2.7-100 times faster than competing state-of-the-art tools on the largest database while being equally accurate. In particular, raxtax exhibits increasing speedups with growing query and reference sequence numbers compared to existing tools (for 100 000 and 1 000 000 query and reference sequences overall, it is 1.3 and 2.9 times faster, respectively), and therefore alleviates the taxonomic classification scalability challenge. AVAILABILITY AND IMPLEMENTATION: raxtax is available at https://github.com/noahares/raxtax under a CC-NC-BY-SA license. The scripts and summary metrics used in our analyses are available at https://github.com/noahares/raxtax_paper_scripts. The source code, sequence data, and summarized results of the analyses are available at https://doi.org/10.5281/zenodo.15057027.

Software

GRUMB: a genome-resolved metagenomic framework for monitoring urban microbiomes and diagnosing pathogen risk.

SUMMARY: Urban infrastructure hosts dynamic microbial communities that complicate biosurveillance and AMR monitoring. Existing tools rarely combine genome-resolved reconstruction with ecological modeling and batch-aware analytics tailored to infrastructure-scale studies. We present GRUMB (Genome-Resolved Urban Microbiome Biosurveillance), an open-source, SLURM-compatible pipeline that reconstructs high-quality metagenome-assembled genomes (MAGs) from shotgun sequencing reads and integrates taxonomic/functional annotation (CARD, VFDB), batch-aware normalization, ecological diagnostics and machine learning classification of environment types with uncertainty and risk scoring. GRUMB accepts either SRA project accessions or paired-end FASTQ files with metadata, and produces assemblies, MAGs, taxonomic and functional profiles, ecological outputs and risk-informed classification. Its modular design enables reproducible, infrastructure-scale biosurveillance across diverse environments. AVAILABILITY AND IMPLEMENTATION: GRUMB is freely available under the MIT License at: https://github.com/SuleimanAminu/genome-resolved-urban-microbiome-biosurveillance; Zenodo DOI: https://doi.org/10.5281/zenodo.15505402. Requirements: Linux (Ubuntu 20.04+), Python 3.11, R 4.2+, SLURM. Issues and feature requests are tracked on GitHub.

Microbiota

[Sound localization in patients with asymmetrical hearing loss].

Good directional hearing ability demands good and symmetrical hearing in both ears. We report the effect of impaired hearing on the directional hearing ability of 98 patients, especially of patients with bilateral asymmetrical hearing loss. The directional testing device included 12 loudspeakers placed at 30 degree intervals in a circle with a diameter of 3.25 m, whose centre lay between the ears of the patient. In included an audiometer for producing the signals, an amplifier and a PDP11/23 computer interfaced to a loudspeaker switch bank. The subject's answers to 60 directionally randomized stimuli were recorded. During the presentation of the signal the patients were not allowed to turn their head. The patients had to name the number of the loudspeaker on the circle that they thought was producing the sound. In addition to the directional hearing test a pure-tone audiogram was done, and the middle- and high-frequency hearing loss estimated. The records of the directional hearing test were analysed in two new ways: firstly, vector analysis of the errors; secondly, averaging of the difference between the true interaural time delay and the virtual time difference, which was implicated in the possibly incorrect answer of the patient (effective delta-t-parameter). This average gives a score for the uncertainty in defining the correct "cone of confusion". In addition to the statistical analysis, two cases are reported showing the directional hearing ability of two patients with neuromas treated by transtemporal surgery, with some residual hearing.(ABSTRACT TRUNCATED AT 250 WORDS)

Audiometry

Characterizing the Uncertainty, Misclassification and Inconsistency of Polygenic Prediction.

Polygenic risk scores (PRSs) hold promise for precision medicine, yet their clinical translation is hindered by substantial uncertainty in individual risk estimates and often limited agreement in risk stratification across multiple PRSs for the same disease. We develop a unified inferential framework to calibrate PRS point estimates and uncertainties for both quantitative traits and binary phenotypes, and to characterize how PRS accuracy, uncertainty, pairwise correlation jointly determine misclassification and classification inconsistency. We show, both theoretically and empirically, that individual- and population-level misclassification and inconsistency rates are highly predictable in independent datasets. We further evaluate PRS integration and uncertainty-aware probabilistic thresholding strategies that reduce misclassification and improve concordance in risk stratification. Together, these results demonstrate that instability in PRS-based classification is a predictable statistical consequence of uncertainty and establish a principled foundation for incorporating uncertainty into PRS-based risk interpretation, communication, and clinical decision-making.

Journal Article

Psychological androgyny and preference in loss-framed gambles of medical students: possible implications for resource utilization.

Physician decisions concerning allocation of health care resources to patients are highly variable and poorly understood. Psychological androgyny theory (PAT) has been employed as a model of the interpersonal and task activities required of physicians for care of their patients. Several studies have successfully predicted physician resource utilization using measures derived from PAT. Using a sample of 97 first-year medical students, the authors explored the relationship between PAT and risk preference in loss-framed gambles in order to elucidate the process whereby variables derived from PAT predict resource utilization. As hypothesized, students selecting the certain loss had significantly higher mean androgyny scores than did students selecting uncertainty. Research involving these constructs is integrated in the context of a theoretical "causal model," which highlights issues deserving of future research.

Choice Behavior

Control-related cognitions and depression among inpatient children and adolescents.

In previous studies, children with numerous depressive symptoms have shown two patterns of control-related cognition: (1) low levels of perceived personal competence, and (2) "contingency uncertainty"--confusion regarding the causes of significant events. The generality of these findings was tested for more seriously disturbed children. Three child inpatient samples, from separate psychiatric hospitals, completed the Children's Depression Inventory (CDI) plus measures of control-related beliefs. In all three samples, the findings resembled those of previous studies: CDI scores were significantly related to low perceived competence and to contingency uncertainty; by contrast, CDI scores were only weakly related to perceived noncontingency. The findings suggest that depressive symptoms in children may be (1) more closely linked to "personal helplessness" than to "universal helplessness," and (2) more closely linked to uncertainty about the causes of events than to firm beliefs in noncontingency. The findings carry implications for etiology and treatment of child depression.

Adolescent

Diagnosis of birth asphyxia on the basis of fetal pH, Apgar score, and newborn cerebral dysfunction.

Imprecise diagnosis of birth asphyxia coupled with uncertainties about causal factors for neurologic abnormalities in the newborn have greatly fueled the current litigation crisis in obstetrics. Our goal was to more precisely define birth asphyxia based on fetal condition as measured by umbilical artery blood pH, Apgar scores, and neurologic condition of newborns. We selected for study 2738 patients with singleton pregnancies with cephalic presentations who were delivered of infants at term to avoid complications such as prematurity, which may affect infant outcome independent of birth condition. The basis for study of these particular patients were defined criteria for high risk and an indicated arterial cord pH value. A total of five infants demonstrated cerebral dysfunction as evidenced by seizures during the neonatal period. Infection was linked to seizures in three of these infants; one infant had neonatal asphyxia and only one infant's clinical course could be attributed solely to birth events (uterine rupture). Stratification of umbilical artery blood pH values, Apgar scores, and combinations of these dependent variables in relation to newborn clinical outcomes revealed that infants must be severely depressed at delivery before birth asphyxia can be reliably diagnosed. Such depression includes Apgar scores less than or equal to 3 at 1 and 5 minutes plus umbilical artery pH values less than 7.00.

Acid-Base Equilibrium

Information retrieval for teaching files: a preliminary study.

A computer algorithm for information retrieval from an electronic teaching file has been developed. This index enables the user to retrieve cases from a teaching file, based on the input of a combination of features. The algorithm is based on nearest neighbor analysis, and is programmed in the "C" language. A teaching file with this index is very easy to use as a reference resource for diagnosing unknown cases. A model was developed for a preliminary test of how likely a user would be to review a teaching file case that is the same diagnosis as an unknown case, thereby reducing uncertainty of diagnosis. The model used 110 cases of arthritis radiographs of hands scored by a skeletal radiologist. The result of the model suggests that the correct diagnosis would be reviewed 83% of the time. A standard method of reducing uncertainty of diagnosis (the maximum likelihood discriminant function) would have picked the correct diagnosis 78% of the time. The results indicate that a teaching file with the computer index is a practical tool for dealing with the uncertainty in diagnosis of unknown cases. The computer index could be included with videodisc-based teaching files (such as the American College of Radiology files). Using teaching files as a reference for interpreting unknown cases may reduce interobserver variability.

Algorithms

Calibrated Prediction Intervals for Polygenic Scores: Updated Comparisons, Contextual Calibration, and Data Normalization.

Calibrated prediction intervals for polygenic scores (PGS) are essential for communicating individual-level uncertainty in genomic medicine. We present updated comparisons of two methods for constructing such intervals: CalPred, a parametric approach, and PredInterval, a non-parametric approach. Our results show that both methods can achieve calibrated coverage, although CalPred additionally requires a sufficiently large calibration set. The two methods also exhibit complementary trade-offs with respect to dataset size and risk identification. We further show that contextual calibration, as introduced in Hou et al. and followed in Shi et al., is most naturally achieved through appropriate phenotype normalization and data preprocessing. Apparent miscalibration can arise from inadequate normalization or from providing contextual information to some methods but not others. In UK Biobank, standard GWAS phenotype normalization procedures are sufficient to achieve contextual calibration for traits analyzed. In the extreme simulations of Hou et al. and Shi et al., supplying contextual covariates to PredInterval restores contextual calibration without normalization, and appropriate normalization can achieve contextual calibration without supplying covariates, while also substantially improving upstream tasks including association power and PGS accuracy. Together, these results underscore the central role of phenotype normalization and data preprocessing in GWAS analyses, including reliable uncertainty quantification for PGS.

Journal Article

Improving the reliability of polygenic risk score-based prediction for cardiovascular and renal complications across ancestries in type 2 diabetes using Mondrian Cross-Conformal Prediction.

Polygenic risk scores (PRS) developed in European populations often show reduced predictive performance in non-European populations, limiting their clinical utility. This lack of transferability across ancestries remains a major challenge in genomic medicine and raises concerns about health equity. We aimed to evaluate whether uncertainty-aware prediction, implemented through Mondrian Cross-Conformal Prediction, improves the performance and reliability of polygenic risk score-based predictions across ancestries for nephropathy, stroke, and myocardial infarction in individuals with type 2 diabetes in a multi-ethnic cohort. We leveraged Mondrian Cross-Conformal Prediction (MCCP), an uncertainty quantification framework, combined with logistic regression applied to a multi-polygenic risk score (multiPRS) to predict the risk of nephropathy, stroke, and myocardial infarction in individuals with type 2 diabetes. Two training frameworks were evaluated: one using 4,098 individuals with type 2 diabetes of European ancestry from the ADVANCE trial for training and 17,574 White British, 1,145 South Asian, and 749 African UK Biobank participants for testing; and another using the 17,574 White British UK Biobank participants for training and the South Asian and African participants for testing. Logistic regression provided robust baseline performance across populations. On top of this baseline, MCCP did not improve performance but added capabilities absent from probability-based stratification: for each individual, it issued a prediction together with an explicit confidence and credibility level; it allowed a tolerated error level to be set in advance and delivered prediction sets respecting it in the majority of settings; and it flagged individuals for whom no reliable prediction could be made. Applying MCCP to PRS-based prediction thus enables uncertainty-aware risk stratification and improves the reliability of risk prediction across ancestries, providing a more equitable framework for clinical use.

Female

Dairy products and the risk of prostatic cancer.

Dietary indicators of prostatic cancer risk were analyzed in a case-control study conducted in Northern Italy on 96 histologically confirmed cases and 292 controls in hospital for acute, nonneoplastic or genital tract diseases. There was a significant trend in risk as regards frequency of milk consumption: compared with nondrinkers or occasional milk drinkers, the relative risk (RR) was 1.2 (95% confidence interval, Cl, 0.7-1.9) for 1 or 2 glasses per day and 5.0 (95% Cl 1.5-16.6) for 2 or more glasses per day. By contrast, no consistent association was observed with measures of cheese or butter intake. This might, at least in part, be attributable to the lower measurement errors for milk (which tends to be consumed in regular and uniform patterns) as compared with other dairy products. However, the interpretation of these findings is not clear, since other sources of animal fat, like eggs or meat, as well as a summary fat score, were unrelated to prostatic cancer. Although these limitations and uncertainties are substantial, this study provides further evidence that elevated milk consumption may be an indicator of prostatic cancer risk.

Aged

The principle of parsimony: Glasgow Coma Scale score predicts mortality as well as the APACHE II score for stroke patients.

Although the development and use of severity-of-illness measures has gained widespread enthusiasm, uncertainty remains as to the optimal measure for stroke patients. The Health Care Financing Administration recently derived a severity-of-illness measure based on the APACHE II system to explain differences in Medicare mortality rates among hospitals treating stroke patients. We hypothesized that the Glasgow Coma Scale score provides prognostic information of accuracy comparable to that of the APACHE II score for stroke patients, yet is simpler and cheaper to abstract from the medical record. We therefore studied 246 patients hospitalized with stroke, including 49 oversampled mortalities. The Glasgow Coma Scale score was as accurate as the APACHE II score in predicting stroke mortality both before (r = -0.50 and r = 0.50, respectively) and after (r = -0.40 and r = 0.38, respectively) the oversampled mortalities were excluded. The APACHE II score required abstraction of 16 variables from the medical record compared with three for the Glasgow Coma Scale score and required more than three times the time to abstract from the medical record. Therefore, in the interest of parsimonious data collection, the Glasgow Coma Scale may be a preferable severity-of-illness measure for patients with stroke.

Aged

The impact of phenotypic variation on genetic analysis: application to X-linkage in manic-depressive illness.

Genetic linkage studies have opened new vistas for behavioral and psychiatric genetics. However, phenotypic diversity and diagnostic uncertainties can lead to spurious linkage findings. A method of analysis is proposed that takes these factors into account. When applied to manic-depressive disease, the results indicate that previous evidence for a major gene localized on the distal long arm of the X-chromosome cannot be ascribed to phenotypic uncertainties and misclassifications, i.e., a type I error. Although the lod score (the logarithm of odds) favoring linkage is reduced with the more restrictive clinical definitions of the phenotype, it remains significant nonetheless. Thus, the linkage finding is robust over a range of phenotypic patterns and presumed phenocopy frequencies. The results also suggest that the X-linked phenotype is a particularly severe form of manic depression characterized by early onset, high familial prevalence of the bipolar form, and high recurrence rate of major depression. These findings may have important implications for the design and interpretation of genetic linkage studies and for refining diagnostic techniques in mental disorders.

Adolescent

Effects of uncertainty on melodic information processing.

In three experiments, musically trained and untrained adults listened to three repetitions of a 5-note melodic sequence followed by a final melody with either the same tune as those preceding it or differing in one position by one semitone. In Experiment 1, ability to recognize the final sequence was examined as a function of redundancy at the levels of musical structure in a sequence, contour complexity of transpositions in a trial, and trial context in a session. Within a sequence, tones were related as the major or augmented triad; within a trial, the four sequences began on successively higher notes (simple macrocontour) or on randomly selected notes (complex macrocontour); and within a session, trials were either blocked (all major or all augmented) or mixed (major and augmented randomly selected). Performance was superior for major melodies, for systematic transpositions within a trial (simple macrocontours), for blocked trials, and for musically trained listeners. In Experiment 2, we examined further the effect of macrocontour. Performance on simple macrocontours exceeded that on complex, and excluded the possibility that repetition of the 20-note sequences provided the entire benefit of systematic transposition in Experiment 1. The effect of musical structure (major/augmented) was also replicated. In Experiment 3, listeners provided structure ratings of ascending 20-note sequences from Experiment 2. Ratings on same trials were higher than those on corresponding different trials, in contrast to performance scores for augmented same and different trials in previous experiments. The concept of functional uncertainty was proposed to account for recognition difficulties on augmented same trials. The significant effects of redundancy on all the levels examined confirm the utility of the information-processing framework for the study of melodic sequence perception.

Adolescent

scRNA-seq and bulk RNA-seq reveal the characteristics of macrophage copper metabolism and establish a risk signature in hepatocellular carcinoma.

BACKGROUND: Hepatocellular carcinoma (HCC) is a prevalent malignancy with an urgent need for improved prognostic stratification and treatment-response prediction. This study aimed to explore a macrophage copper metabolism-associated prognostic model and to investigate the relationship between this risk model and the tumor immune microenvironment. METHODS: The FindClusters function was used to analyze cell clusters, and CellChat and CellPhoneDB/LIANA were employed for cell-cell communication analysis. Copper metabolism-related genes were sourced from the MSigDB database. A prognostic risk model was established using least absolute shrinkage and selection operator (LASSO) analysis and multivariate Cox regression analysis, and a nomogram was constructed by integrating the prognostic model with clinicopathological factors. Additional analyses were performed to map the seven model genes in single-cell data, assess model uncertainty and robustness, evaluate macrophage/copper/cuproptosis-related transcriptional programs, and examine the correlations between risk score, immune infiltration and predicted drug sensitivity. RESULTS: Using single-cell RNA sequencing (scRNA-seq) data, we identified four macrophage subpopulations. Macrophages with high SPP1 expression showed close interaction with T cell populations and were associated with copper ion metabolism. By incorporating 141 copper metabolism-related genes and using The Cancer Genome Atlas Liver Hepatocellular Carcinoma (TCGA-LIHC) cohort, we constructed a seven-gene risk prediction model. Additional single-cell mapping showed that the model genes were detectable in the HCC single-cell dataset and showed a macrophage-associated expression pattern. The model showed moderate prognostic discrimination in TCGA-LIHC, whereas its external performance was heterogeneous and remained evaluable across external cohorts, with performance varying among datasets. Immune and mechanism-related analyses suggested that the risk signature was associated with macrophage-related infiltration, copper metabolism and cuproptosis-related transcriptional programs. Drug sensitivity analysis nominated Daporinad as a computationally predicted candidate compound, supporting Daporinad as a pharmacogenomic candidate for follow-up investigation. CONCLUSIONS: By integrating scRNA-seq and bulk RNA sequencing (RNA-seq) data, we constructed a macrophage copper metabolism-associated prognostic signature for HCC. The risk score was associated with survival, immune microenvironment features and predicted drug response, providing a transcriptomic framework for risk stratification and therapeutic hypothesis generation.

Hepatocellular carcinoma (HCC)

Population-scale detection of methylation outliers from long-read genome sequencing.

BACKGROUND: Aberrant DNA methylation can mediate the functional effects of rare genetic variation and contribute to imprinting disorders, repeat expansion diseases, and other pathogenic regulatory mechanisms. Long-read sequencing technologies now enable genome-wide detection of CpG methylation alongside genetic variation from a single assay. However, methods for systematic identification and interpretation of methylation outliers from long-read sequencing data remain limited. METHODS: We developed METAFORA, a computational workflow for detecting methylation outlier regions from PacBio and Oxford Nanopore long-read sequencing data. METAFORA constructs population-level methylation references, segments the genome into correlated CpG blocks, infers technical and biological sources of variation through hidden factor estimation, models uncertainty due to variable depth sequencing, and computes covariate-adjusted methylation outlier scores for individual samples. We applied METAFORA across large long-read sequencing cohorts and integrated methylation outliers with multi-omic data. METAFORA is implemented as a snakemake workflow available at https://github.com/tjense25/METAFORA. RESULTS: METAFORA identified methylation outlier regions associated with rare structural variants, tandem repeat expansions, and imprinting abnormalities. We found outlier regions were enriched for molecular outliers across transcriptomic and chromatin accessibility datasets, supporting their functional relevance in gene regulation. In a representative case, METAFORA identified an imprinting defect affecting the GNAS locus associated with an STX16 deletion. CONCLUSIONS: METAFORA enables scalable detection and interpretation of methylation outliers from long-read sequencing data and provides a framework for integrating epigenetic outliers with genomic and multi-omic analyses. These approaches may improve interpretation of rare regulatory variation and support discovery of clinically relevant epigenetic abnormalities in genomic medicine.

DNA methylation