Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,045 records · Page 58Linked to original sources

STRUMP-I: Structure-based machine learning approach to pMHC-I binding prediction using force field energy features.

The adaptive immune system monitors cellular integrity by recognizing short peptides from intracellular proteins presented on Major Histocompatibility Complex class I (MHC-I) molecules, collectively termed peptide-MHC complexes (pMHC), enabling detection of foreign or mutated proteins. With the rising importance of immunotherapies targeting neoantigens in cancers, the ability to accurately predict which peptides will bind to the diverse population of MHC alleles is critically important. Current computational methods for pMHC-I prediction fall broadly into sequence-based methods, which rely heavily on large training datasets, and structure-based methods that leverage structural modeling and energetics of pMHC binding. While sequence-based methods have been popularly used, their performance is dependent on the size and quality of training data. On the other hands, while structure-based approaches can generalize better across diverse MHC alleles, they traditionally depend on identifying a single global minimum energy conformation, an assumption that often fails due to the inherent binding promiscuity of MHC-I molecules. To address these limitations, we developed a STRUMP-I (STRUcture-based pMHC Prediction (for class I)), a novel pMHC binding prediction tool that directly leverages a broad set of force-field-derived energy terms as machine-learning features. STRUMP-I achieves performance comparable to state-of-the-art sequence-based models while significantly outperforming them on MHC alleles with limited representation in training data. Furthermore, STRUMP-I demonstrates strong synergy when integrated with sequence-based methods, notably enhancing prediction precision. The robustness and generalizability of STRUMP-I were confirmed by evaluating its predictive performance on independent, previously unseen datasets, including an experimentally validated cancer neoantigen dataset. This combined approach advances our capability to reliably identify clinically relevant neoantigen targets. The source code and trained models are available at https://github.com/yoonjoolab/STRUMP-I.

energy optimization↗

Systematic evaluation of metatranscriptomic differential gene expression in silico, in vitro, and in vivo enables elucidation of inter-species cross-feeding.

Metatranscriptomic (MTX) sequencing quantifies gene expression from the collective genomes of microbial communities (microbiomes), enabling assessment of functional activity rather than functional potential. While differential expression testing is instrumental to RNA-sequencing analysis, current metatranscriptomic approaches have been benchmarked only on simulated data and not under real operating conditions, resulting in a lack of standard practices. Here, we evaluate the performance of statistical differential expression methods on both simulated datasets and data collected from real bacterial 'mock communities' designed for this purpose. We assess the robustness of individual methods to organisms' low relative abundance, differential abundance, low prevalence, and transcription rate changes, showing that no existing methods perform adequately across all confounding conditions. We then apply the same approaches to metatranscriptomic datasets generated from gnotobiotic mice colonized with defined consortia of human bacterial strains and show that the method nominated by our mock community comparisons successfully inferred cross-feeding dynamics which were validated in vitro. We conclude that MTX method benchmarking on real, not simulated, datasets can and should optimize model implementation, enabling inference and validation of cross-feeding and other inter-species and host-microbe dynamics from in vivo studies.

Journal Article↗

The Annotated Blueprint: Integrated Functional Genomic Resources for a model Tetraploid Wheat Triticum turgidum cv. Kronos.

Triticum turgidum cv. Kronos is a tetraploid wheat cultivar that underpins one of the richest community platforms for functional genomics. Over the past decade, about 3,000 exome- and promoter-capture datasets, linked to mutagenized seed stocks, and transcriptomic and phenotypic resources have accumulated, yet the absence of a reference genome has constrained their impact. Here, we present a chromosome-scale reference genome of Kronos with high-confidence annotations, including manual curation of over 1,000 disease resistance (NLR) genes. This reference revealed previously hidden NLR diversity and clarified their genomic organization at chromosomal ends. Re-analysis of exome- and promoter-capture datasets enabled high-resolution mutation discovery in genes and regulatory regions that were previously inaccessible, uncovering the full standing variation present in Kronos mutant lines. We further re-curated transcriptomic and small RNA datasets, generating improved, genome-wide maps of microRNAs and phasiRNAs important for wheat development. Collectively, these resources elevate Kronos to reference quality and establish it as a versatile platform for functional and translational wheat research.

Journal Article↗

Elastic registration of fMRI data using Bézier-spline transformations.

A three-dimensional (3-D) elastic registration algorithm has been developed to find a veridical transformation that maps activation patterns from functional magnetic resonance imaging (fMRI) experiments onto a 3-D high-resolution anatomical dataset. The proposed algorithm uses trilinear Bézier-splines and a 3-D voxel-based optimization technique to determine the transformation that maps the functional data onto the coordinate system of the anatomical dataset. Simple conditions are presented which guarantee that the data are mapped one-to-one on each other. Two voxel-based similarity measures, the linear correlation coefficient and the entropy correlation coefficient, are used. Their performance with respect to the registration of fMRI data is compared. Tests on simulated and real data have been performed to evaluate the accuracy of the method. Our results demonstrate that subvoxel accuracy can be achieved even for noisy low-resolution multislice datasets with local distortions up to 10 mm. Although the method is optimized for the registration of functional and anatomical MR images, it can also be used for solving other elastic registration problems.

Algorithms↗

Dynamic HIV/AIDS parameter estimation with application to a vaccine readiness study in southern Africa.

This paper proposes a procedure of parameter estimation for all parameters of the three-dimensional HIV model. The least square based procedure uses standard optimization routines to allow parameter extraction for individual patients. It is shown how additional information from outside a measurement dataset can be included in the estimation routine to increase the reliability and accuracy of parameter estimates. A dataset from 44 patients of Southern Africa is analyzed to find the set point and the time until set point for these patients together with an estimate of the model parameters with confidence intervals for the cohort. The procedure is also applied to a long-term dataset of the HIV/AIDS progression to find possible variations in parameters.

Acquired Immunodeficiency Syndrome↗

Maximizing sensitivity in medical diagnosis using biased minimax probability machine.

The challenging task of medical diagnosis based on machine learning techniques requires an inherent bias, i.e., the diagnosis should favor the "ill" class over the "healthy" class, since misdiagnosing a patient as a healthy person may delay the therapy and aggravate the illness. Therefore, the objective in this task is not to improve the overall accuracy of the classification, but to focus on improving the sensitivity (the accuracy of the "ill" class) while maintaining an acceptable specificity (the accuracy of the "healthy" class). Some current methods adopt roundabout ways to impose a certain bias toward the important class, i.e., they try to utilize some intermediate factors to influence the classification. However, it remains uncertain whether these methods can improve the classification performance systematically. In this paper, by engaging a novel learning tool, the biased minimax probability machine (BMPM), we deal with the issue in a more elegant way and directly achieve the objective of appropriate medical diagnosis. More specifically, the BMPM directly controls the worst case accuracies to incorporate a bias toward the "ill" class. Moreover, in a distribution-free way, the BMPM derives the decision rule in such a way as to maximize the worst case sensitivity while maintaining an acceptable worst case specificity. By directly controlling the accuracies, the BMPM provides a more rigorous way to handle medical diagnosis; by deriving a distribution-free decision rule, the BMPM distinguishes itself from a large family of classifiers, namely, the generative classifiers, where an assumption on the data distribution is necessary. We evaluate the performance of the model and compare it with three traditional classifiers: the k-nearest neighbor, the naive Bayesian, and the C4.5. The test results on two medical datasets, the breast-cancer dataset and the heart disease dataset, show that the BMPM outperforms the other three models.

Algorithms↗

A support vector machines classifier to assess the severity of idiopathic scoliosis from surface topography.

A support vector machines (SVM) classifier was used to assess the severity of idiopathic scoliosis (IS) based on surface topographic images of human backs. Scoliosis is a condition that involves abnormal lateral curvature and rotation of the spine that usually causes noticeable trunk deformities. Based on the hypothesis that combining surface topography and clinical data using a SVM would produce better assessment results, we conducted a study using a dataset of 111 IS patients. Twelve surface and clinical indicators were obtained for each patient. The result of testing on the dataset showed that the system achieved 69-85% accuracy in testing. It outperformed a linear discriminant function classifier and a decision tree classifier on the dataset.

Adolescent↗

Simultaneous registration and activation detection for fMRI.

Registration using the least-squares cost function is sensitive to the intensity fluctuations caused by the blood oxygen level dependent (BOLD) signal in functional MRI (fMRI) experiments, resulting in stimulus-correlated motion errors. These errors are severe enough to cause false-positive clusters in the activation maps of datasets acquired from 3T scanners. This paper presents a new approach to resolving the coupling between registration and activation. Instead of treating the two problems as individual steps in a sequence, they are combined into a single least-squares problem and are solved simultaneously. Robustness tests on a variety of simulated three-dimensional EPI datasets show that the stimulus-correlated motion errors are removed, resulting in a substantial decrease in false-positive and false-negative activation rates. The new method is also shown to decorrelate the motion estimates from the stimulus by testing it on different in vivo fMRI datasets acquired from two different 3T scanners.

Algorithms↗

Characterization of spiculation on ultrasound lesions.

Spiculation is a stellate distortion caused by the intrusion of breast cancer into surrounding tissue. Its existence is an important clue to characterizing malignant tumors. Many successful mammographic methods have been proposed to detect tumors with spiculation. Traditional two-dimensional (2-D) ultrasound cannot easily find spiculations because spiculations normally appear parallel to the surface of the skin. Recently, three-dimensional (3-D) ultrasound has been gradually used in clinical applications and it has been proven to be useful in determining the architectural distortion or spiculation that surrounds a breast tumor. This paper aims to identify spiculation from 3-D ultrasonic volume data of a tumor found by a physician. In the proposed method, each coronal slice of volume data is successively extracted and then analyzed as a 2-D ultrasound image by the proposed spiculation detection method. First, in each horizontal slice, the modified rotating structuring element (ROSE) operation is used to find the central region in which spiculation lines converge. Second, the stick algorithm is used to estimate the direction of the edge of each pixel around the central region. A pixel whose edge points toward the central region is marked as a potential spiculation. Finally, the marked pixels are collected around the central region and their distribution is analyzed to determine whether spiculation is present. The 3-D test datasets were obtained using the Voluson 530 or 730, Kretztechnik, Austria. First, the proposed method was tested on 104 2-D typical coronal images (selected by an experienced physician) extracted from 52 3-D ultrasonic datasets. Finally, 225 3-D pathologically proven datasets were tested to evaluate the performance. Spiculations are more easily observed in the coronal view than in the other two views. That is, the 3-D ultrasound is a powerful tool for identifying spiculations. Furthermore, 16% (19/120) of benign cases and 90% (94/105) of malignant cases are detected as spiculations.

Algorithms↗

Study of temporal stationarity and spatial consistency of fMRI noise using independent component analysis.

Spatial independent component analysis (ICA) was used to study the temporal stationarity and spatial consistency of structured functional MRI (fMRI) noise. Spatial correlations have been used in the past to generate filters for the removal of structured noise for each time-course in an fMRI dataset. It would be beneficial to produce a multivariate filter based on the same principles. ICA is examined to determine if it has properties that are beneficial for this type of filtering. Six fMRI baseline datasets were decomposed via spatial ICA. The time-courses associated with each component were tested for wide-sense stationarity using the wide sense stationarity quotient (WSS). Each dataset was divided into three subsets and each subset was decomposed. The components of first and third subset were matched by the strength of their correlation. The components produced by ICA were found to have largely nonstationary time-courses. Despite the temporal nonstationarity in the data, ICA was found to produce consistent spatial components. The degree of correlation among components differed depending on the amount of dimension reduction performed on the data. It was found that a relatively small number of dimensions produced components that are potentially useful for generating a spatial fMRI filter.

Algorithms↗

Improved K-means clustering algorithm for exploring local protein sequence motifs representing common structural property.

Information about local protein sequence motifs is very important to the analysis of biologically significant conserved regions of protein sequences. These conserved regions can potentially determine the diverse conformation and activities of proteins. In this work, recurring sequence motifs of proteins are explored with an improved K-means clustering algorithm on a new dataset. The structural similarity of these recurring sequence clusters to produce sequence motifs is studied in order to evaluate the relationship between sequence motifs and their structures. To the best of our knowledge, the dataset used by our research is the most updated dataset among similar studies for sequence motifs. A new greedy initialization method for the K-means algorithm is proposed to improve traditional K-means clustering techniques. The new initialization method tries to choose suitable initial points, which are well separated and have the potential to form high-quality clusters. Our experiments indicate that the improved K-means algorithm satisfactorily increases the percentage of sequence segments belonging to clusters with high structural similarity. Careful comparison of sequence motifs obtained by the improved and traditional algorithms also suggests that the improved K-means clustering algorithm may discover some relatively weak and subtle sequence motifs, which are undetectable by the traditional K-means algorithms. Many biochemical tests reported in the literature show that these sequence motifs are biologically meaningful. Experimental results also indicate that the improved K-means algorithm generates more detailed sequence motifs representing common structures than previous research. Furthermore, these motifs are universally conserved sequence patterns across protein families, overcoming some weak points of other popular sequence motifs. The satisfactory result of the experiment suggests that this new K-means algorithm may be applied to other areas of bioinformatics research in order to explore the underlying relationships between data samples more effectively.

Algorithms↗

A data analysis competition to evaluate machine learning algorithms for use in brain-computer interfaces.

We present three datasets that were used to conduct an open competition for evaluating the performance of various machine-learning algorithms used in brain-computer interfaces. The datasets were collected for tasks that included: 1) detecting explicit left/right (L/R) button press; 2) predicting imagined L/R button press; and 3) vertical cursor control. A total of ten entries were submitted to the competition, with winning results reported for two of the three datasets.

Algorithms↗

A phase space spline smoother for fitting trajectories.

This paper presents a phase space spline smoother, which is especially useful for finding a best-fit trajectory from multiple examples of a given physical motion. Unlike conventional spline smoothers, the phase space spline smoother can simultaneously fit position and velocity information. The use of velocity information is important for modeling the dynamic motion of physical systems because the state space of these systems typically includes both position and velocity variables. A detailed description of the computational procedure is presented, along with a discussion of computational expense and practical guidelines for variance estimation, preprocessing of the target dataset, smoothing of multidimensional datasets, and cross-validation for selection of smoothing weights. The smoother is demonstrated on a dataset of handwriting motions.

Journal Article↗

Air-coupled through-transmission fan-beam tomography using divergent capacitive ultrasonic transducers.

Abstracttrasonic transducers (CUTs) with curved backplates was used to acquire signals through regions of air containing solid objects, air flow, and temperature fields. Fan-beam datasets were collected and used in a tomographic reconstruction algorithm to produce cross-sectional images of the area under interrogation. In the case of the solid objects, occluded rays from the projections were accounted for using a compensation algorithm and a priori knowledge of the object. A rebinning routine was used to pick out parallel ray sets from the fan-beam data. The effects of further reducing the number of datasets also were investigated, and, in the case of imaging solid objects, characteristic Gibbs phenomena were seen in the reconstructions as expected. However, when imaging temperature and flow fields, the aliasing artefacts were not seen, but the reconstructed values decreased with the size of dataset used. The effect of changing the kernel filter function also was investigated, with the different filters giving the best compromise between image noise, reconstruction accuracy, and amount of data required in each scenario.

Algorithms↗

Benchmarking of Reference-Based Tools for Strain-Level Resolution of Plant Microbiome.

Strain-level identification of each microbe is crucial for understanding its role in the host. Most of the existing tools have primarily been evaluated on human metagenomic datasets, whereas the plant microbiome exhibits greater diversity and complexity and thus poses a challenge in the strain-level resolution of individual microbes. In this study, we conducted a comprehensive benchmarking of available reference-based tools for strain-level resolution of the plant microbiome. We evaluated seven tools on various performance parameters, like computational requirements, F1-score and relative abundances using synthetic datasets comprising microbes known to have strong associations with plants as well as real plant microbiome datasets. Our results demonstrated a better performance of StrainScan on the synthetic data, achieving higher F1-score and more accurate relative abundance estimates as compared to other tools, but its performance declined gradually with increasing strain diversity. However, StrainGE and StrainScan exhibited competitive performance on real plant metagenome data. Overall, though StrainGE exhibited better performance, it was more computationally expensive. However, StrainScan performed better in detecting low-abundance strains. Our findings suggest the comparative suitability of the available tools for the strain-level analysis of plant metagenome data and highlight the need for the development of more efficient and accurate taxonomic classifiers capable of handling the complex plant metagenome data while maintaining computational efficiency.

Microbiota↗

Interpretation of the results of common principal components analyses.

Common principal components (CPC) analysis is a new tool for the comparison of phenotypic and genetic variance-covariance matrices. CPC was developed as a method of data summarization, but frequently biologists would like to use the method to detect analogous patterns of trait correlation in multiple populations or species. To investigate the properties of CPC, we simulated data that reflect a set of causal factors. The CPC method performs as expected from a statistical point of view, but often gives results that are contrary to biological intuition. In general, CPC tends to underestimate the degree of structure that matrices share. Differences of trait variances and covariances due to a difference in a single causal factor in two otherwise identically structured datasets often cause CPC to declare the two datasets unrelated. Conversely, CPC could identify datasets as having the same structure when causal factors are different. Reordering of vectors before analysis can aid in the detection of patterns. We urge caution in the biological interpretation of CPC analysis results.

Analysis of Variance↗

Contrasting patterns of radiation in African and Australian Restionaceae.

The floras of the Mediterranean-climate areas of southern Africa and southwestern Australia are remarkably species rich. Because the two areas are at similar latitudes and in similar positions on their respective continents, they have probably had similar Cenozoic climatic histories. Here we test the prediction that the evolution of the species richness in the two areas followed a similar temporal progression by comparing the rates of lineage accumulation for African and Australian Restionaceae. Restionaceae (Poales) are typical and often dominant elements in the fynbos vegetation of the Cape Floristic Region of southern Africa and the kwongan vegetation of the Southwestern Floristic Province of Western Australia. The phylogeny of the family was estimated from combined datasets for rbcL and trnL-F sequences and a large morphological dataset; these datasets are largely congruent. The monophyly of Restionaceae is supported and a basal division into an African clade (approximately 350 species) and an Australian clade (146 species) corroborated. There is also support for a futher subdivision of these two large sister-clades, but the terminal resolution within the African clade is very weak. Fossil pollen records provided a minimum age of the common ancestor of Australian and African Restionaceae as 64-71 million years ago, and this date was used to calibrate a molecular clock. A molecular clock was rejected by a likelihood ratio test; therefore, rate changes between the lineages were smoothed using nonparametric rate smoothing. The rate-corrected ages were used to construct a plot of lineages through time. During the Palaeogene the Australian lineage diversity increased consistent with the predictions of the constant birthrate model, while the African lineage diversity showed a dramatic increase in diversification rate in the Miocene. Incomplete sampling obscures the patterns in the Neogene, but extending the trends to the modern extant diversity suggests that this acceleration in the speciation rate continued in the African clade, whereas the Australian clade retained a constant diversification rate. The substantial morphological and anatomical similarity between the African and Australian Restionaceae appear to preclude morphological innovations as possible explanations for the intercontinental differences. Most likely these differences are due to the greater geographical extent and ecological variation in temperate Australia than temperate Africa, which might have provided refugia for basal Restionaceae lineages, whereas the more mountainous terrain of southern Africa might have provided the selective regimes for a more rapid, recent speciation.

Africa↗

Clinical toxicology: clinical science to public health.

1. The aims of the present paper are to: (i) review progress in clinical toxicology over the past 40 years and to place it in the context of modern health care by describing its development; and (ii) illustrate the use of clinical toxicology data from Scotland, in particular, as a tool for informing clinical care and public health policy with respect to drugs. 2. A historical literature review was conducted with amalgamation and comparison of a series of published and unpublished clinical toxicology datasets from NPIS Edinburgh and other sources. 3. Clinical databases within poisons treatment centres offer an important method of collecting data on the clinical effects of drugs in overdose. These data can be used to increase knowledge on drug toxicity mechanisms that inform licensing decisions, contribute to evidence-based care and clinical management. Combination of this material with national morbidity datasets provides another valuable approach that can inform public health prevention strategies. 4. In conclusion, clinical toxicology datasets offer clinical pharmacologists a new study area. Clinical toxicology treatment units and poisons information services offer an important health resource.

Analgesics, Opioid↗