Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Stigmata: an algorithm to determine structural commonalities in diverse datasets.

An algorithm, Stigmata, is described, which extracts structural commonalities from chemical datasets. It is discussed using several illustrative examples and a pharmaceutically interesting set of dopamine D2 agonists. The commonalities are determined using two-dimensional topological chemical descriptions and are incorporated into the key feature of the algorithm, the modal fingerprint. Flexibility is built into the algorithm by means of a user-defined threshold value, which affects the information content of the modal fingerprint. The use of the modal fingerprint as a diversity assessment tool, as a database similarity query, and as a basis for color mapping the determined commonalities back onto the chemical structures is demonstrated.

Algorithms↗

Locating biologically active compounds in medium-sized heterogeneous datasets by topological autocorrelation vectors: dopamine and benzodiazepine agonists.

Electronic properties located on the atoms of a molecule such as partial atomic charges as well as electronegativity and polarizability values are encoded by an autocorrelation vector accounting for the constitution of a molecule. This encoding procedure is able to distinguish between compounds being dopamine agonists and those being benzodiazepine receptor agonists even after projection into a two-dimensional self-organizing network. The two types of compounds can still be distinguished if they are buried in a dataset of 8323 compounds of a chemical supplier catalog comprising a wide structural variety. The maps obtained by this sequence of events, calculation of empirical physicochemical effects, encoding in a topological autocorrelation vector, and projection by a self-organizing neural network, can thus be used for searching for structural similarity, and, in particular, for finding new lead structures with biological activity.

Databases, Factual↗

Enhanced covariance spectroscopy from minimal datasets.

A novel approach is described for the determination of reliable high-resolution homonuclear NMR covariance spectra from minimal datasets. It uses a sparse sampling scheme along the indirect dimension together with a comprehensive analysis of finite sampling effects that eliminates spurious correlations. The scheme, which is demonstrated for TOCSY and COSY, offers a substantial speed up over current methods, rendering it suitable for high-throughput screening applications.

Image Enhancement↗

Ligand intramolecular motions in ligand-protein interaction: ALPHA, a novel dynamic descriptor and a QSAR study with extended steroid benchmark dataset.

The role of intramolecular motions in ligand-macromolecule interactions has been explored by developing and validating ALPHA, a novel QSAR (quantitative structure-activity relationship) descriptor. It is based on the spectral exponents (alpha), which measure the degree of 1/f alpha noise of coordinate fluctuations in molecular dynamics (MD) simulations. ALPHA is the first truly 'dynamic' QSAR descriptor, i.e., it can be derived directly from an MD trajectory. The performance of ALPHA was tested in detail employing the CBG (corticosteroid binding globulin) affinity of 31 benchmark steroids, supplemented with 11 steroids as an external test set. The only fair (42-50%) correlations of ALPHA with static 3D and electronic descriptors mean that ALPHA forms an independent molecular property. Furthermore, inclusion of ALPHA in the SOMFA/ESP model improves the correlation coefficient from 0.86 to 0.91, and /delta/ave from 0.46 to 0.36 for the benchmark dataset. The predictive ability of ALPHA can be interpreted as indirect evidence of the dynamic contribution to ligand-macromolecule interactions. The physical background of ALPHA is discussed and the importance of molecular motions for biological activity is anticipated.

Ligands↗

Commentary and opinion: III. Some nonontological and functionally unconnected views on current issues in the analysis of PET datasets.

Strother et al. (1995) and Friston (1995) both raise important issues and provide useful reviews of various aspects of PET data analysis. Statisticians would not assume that any single piece of methodology would answer all questions about a type of data in a variety of experimental and observational contexts. The fundamental importance of hypothesis-driven inference, based on well designed experiments, cannot be overestimated for its ability to progress scientific understanding in an orderly manner. However, hypothesis-generating experiments are also vital in their own right. In practice, we generally do not have the luxury of both types of experiment, and we should note Strother et al.'s comment on the importance of extracting as much information as possible from each dataset. Friston (1995) also sees formal testing methods and exploratory methods such as principal components analysis as complementary. The correct approach would therefore seem to be (a) to select methods for formal and exploratory data analysis from the rich existing tool kit of statistical procedures, (b) to modify these as necessary to deal with special PET problems such as multiplicity, (c) to be aware of the assumptions underlying the methods being used and to investigate the problems that can arise if these assumptions fail to hold, (d) to appreciate the complexity of both PET data and of the potential questions that can be asked of it, and (e) to be aware of the limitations of any statistical analysis and the need for caution in interpreting conclusions not based on any predefined hypothesis.

Analysis of Variance↗

Detecting and quantifying circular RNAs in terabyte-scale RNA-seq datasets with CIRI3.

To address recent challenges in circular RNA (circRNA) analysis, we present CIRI3, a tool for circRNA detection and quantification in terabyte-scale RNA-sequencing datasets. Using dynamic multithreaded task partitioning and a blocking search strategy for junction reads, CIRI3 is an order of magnitude faster than existing tools, while providing increased accuracy. We identified differentially spliced circRNAs across 2,535 cancer-related samples, and constructed a pretraining model and a biomarker network provided as the CIRIonco database.

RNA, Circular↗

Automated reconstruction of curvilinear fibres from 3D datasets acquired by X-ray microtomography.

The characterization of fibrous structures is important in both composites and textiles research for relating to the bulk properties of the material. However, the microscopic nature of the fibres and their high densities make them very difficult to characterize. Many techniques have been developed for the measurement and characterization of fibrous structures but they tend to be restricted to measurements on the sample surface or within physical cross-sections. X-ray microtomography can be used to non-destructively probe the internal structure of a range of fibrous materials, providing large amounts of 3D data. A technique has been developed for tracing fibres within 3D datasets acquired by X-ray microtomography and this has been applied to a glass fibre reinforced composite and also a non-woven textile sample. The 3D fibrous structures of both samples were successfully reconstructed and their fibre orientation distributions calculated. This technique enables novel characterizations, such as the through-thickness variation of fibre orientation in non-wovens.

Journal Article↗

The application of new software tools to quantitative protein profiling via isotope-coded affinity tag (ICAT) and tandem mass spectrometry: I. Statistically annotated datasets for peptide sequences and proteins identified via the application of ICAT and tandem mass spectrometry to proteins copurifying with T cell lipid rafts.

Lipid rafts were prepared according to standard protocols from Jurkat T cells stimulated via T cell receptor/CD28 cross-linking and from control (unstimulated) cells. Co-isolating proteins from the control and stimulated cell preparations were labeled with isotopically normal (d0) and heavy (d8) versions of the same isotope-coded affinity tag (ICAT) reagent, respectively. Samples were combined, proteolyzed, and resultant peptides fractionated via cation exchange chromatography. Cysteine-containing (ICAT-labeled) peptides were recovered via the biotin tag component of the ICAT reagents by avidin-affinity chromatography. On-line micro-capillary liquid chromatography tandem mass spectrometry was performed on both avidin-affinity (ICAT-labeled) and flow-through (unlabeled) fractions. Initial peptide sequence identification was by searching recorded tandem mass spectrometry spectra against a human sequence data base using SEQUEST software. New statistical data modeling algorithms were then applied to the SEQUEST search results. These allowed for discrimination between likely "correct" and "incorrect" peptide assignments, and from these the inferred proteins that they collectively represented, by calculating estimated probabilities that each peptide assignment and subsequent protein identification was a member of the "correct" population. For convenience, the resultant lists of peptide sequences assigned and the proteins to which they corresponded were filtered at an arbitrarily set cut-off of 0.5 (i.e. 50% likely to be "correct") and above and compiled into two separate datasets. In total, these data sets contained 7667 individual peptide identifications, which represented 2669 unique peptide sequences, corresponding to 685 proteins and related protein groups.

Amino Acid Sequence↗

A dataset of human fetal liver proteome identified by subcellular fractionation and multiple protein separation and identification technology.

A high throughput process including subcellular fractionation and multiple protein separation and identification technology allowed us to establish the protein expression profile of human fetal liver, which was composed of at least 2,495 distinct proteins and 568 non-isoform groups identified from 64,960 peptides and 24,454 distinct peptides. In addition to the basic protein identification mentioned above, the MS data were used for complementary identification and novel protein mining. By doing the analysis with integrated protein, expressed sequence tag, and genome datasets, 223 proteins and 15 peptides were complementarily identified with high quality MS/MS data.

Cell Membrane↗

Patterns and correlates of treatment: findings of the 2000-2001 NSW minimum dataset of clients of alcohol and other drug treatment services.

The aim of this study was to provide an overview of the first year of the NSW Minimum Dataset for Alcohol and Other Drug Treatment Services data collection, including describing the patterns and correlates of people having received treatment in New South Wales. All closed treatment episodes for the 2000-2001 financial year were included for descriptive, univariate and multivariate analyses. There were 33,459 closed episodes of care in New South Wales in the 2000/2001 financial year. The majority of clients (69%) were male and the mean age was almost 34 years. The majority of treatment is sought for problems related to alcohol (37%) and heroin (33%) use. More than a third (40%) of clients were new to drug and alcohol treatment. Half the clients had a history of injecting drug use with 6.3% of those with heroin as their principal drug of concern, never having injected. The most common main service provided was in-patient withdrawal (26%). Multivariate logistic regression revealed that being older, not homeless, non-indigenous and having heroin as the principal drug of concern predicted receiving out-patient withdrawal management. Analyses of length of stay in residential treatments and number of service contacts in non-residential treatments are reported. The NSW MDS AODTS is a critical information source for policy development, service planning and surveillance. The results of this paper illustrate the utility of the data collection for identifying emerging issues in the patterns of drug use and service delivery for clients with alcohol and other drug problems.

Adolescent↗

A Kinematic Model of the Upper Limb Based on the Visible Human Project (VHP) Image Dataset.

A kinematic model of the arm was developed using high-resolution medical images obtained from the National Library of Medicine's Visible Human Project (VHP) dataset. The model includes seven joints and uses thirteen degrees of freedom to describe the relative movements of seven upper-extremity bones: the clavicle, scapula, humerus, ulna, radius, carpal bones, and hand. Two holonomic constraints were used to model the articulation between the scapula and the thorax. The kinematic structure of each joint was based on anatomical descriptions reported in the literature. The three joints comprising the shoulder girdle - the sternoclavicular joint, the acromioclavicular joint, and the glenohumeral joint - were each modeled as a three degree-of-freedom, ideal, ball-and-socket joint. The articulations at the elbow and wrist - humeroulnar flexion-extension, radioulnar pronation-supination, radiocarpal flexion-extension, and radiocarpal radial-ulnar deviation - were each modeled as a single degree-of-freedom, ideal, hinge joint. Locations of the joint centers and joint axes were derived by graphically inspecting the three-dimensional surfaces of the reconstructed bones. The relative positions of the bones were defined by fixing a reference frame to each bone; the position and orientation of each reference frame were based on the positions of anatomical landmarks and on the shapes of the reconstructed bone surfaces. Tables are provided which specify the positions and orientations of the joint axes and the bone-fixed reference frames for the model arm.

Journal Article↗

Omnibus permutation tests of the overall null hypothesis in datasets with many covariates.

Tests of the overall null hypothesis in datasets with one outcome variable and many covariates can be based on various methods to combine the p-values for univariate tests of association of each covariate with the outcome. The overall p-value is computed by permuting the outcome variable. We discuss the situations in which this approach is useful and provide several examples. We use simulations to investigate seven omnibus test statistics and find that the Anderson-Darling and Fisher's statistics are superior to the others.

Computer Simulation↗

A web management service applied to a comprehensive characterization of Visible Human Dataset colour images.

Visible Human Dataset (VHD) is a remarkable piece of raw digital anatomical knowledge still to be fully exploited. Colours of VHD anatomic images are the natural targets of different algorithmic approaches devoted to understanding the content of the complex digital medical images, but they have never been analysed exhaustively. A full colorimetric characterization of all 9000 VHD colour images may help to take advantage of implicit available information in raw data. This study describes a novel colorimetric characterization and a Visual Knowledge Discovery tool, using methods from database field, data visualization, and image analysis. The applied heterogeneous methods allowed us to develop a histogram meta database and make it available remotely. It consists of a histogram-based colorimetric characterization of the all VHD 24-bit colour images. A user-friendly, interactive, and intuitive 3D framework providing 3D services was built and made freely available. It allows real-time analysis of colour component characteristics of a user-defined set of VHD images providing 3D interactive navigation of the histogram meta database. New knowledge can be discovered using our tool and the histogram meta database provided. This work allowed us to propose novel methods for colour image characterization and obtained results using developed service on VHD colour images let us to partially understand the not fully satisfactorily results achieved so far analysing these images.

Anatomy↗

Versailles minimal dataset for diagnosis of ALS: a distillate of the 2nd Consensus Conference on accelerating the diagnosis of ALS. Versailles 2nd Consensus Conference participants.

The 2nd Consensus Conference (Versailles) recommended that an ALS knowledge-base for initial healthcare providers, diagnosing neurologists and confirming neurologists should be defined to include a simplified version of diagnostic criteria less formal than the World Federation of Neurology El Escorial Revisted Criteria ('ALS diagnosis - An algorithm'), a set of rules concerning red flags which should increase the suspicion of ALS as the diagnosis and minimize the time between suspicion and referral for confirmation of diagnosis ('ALS axioms of referral'), as well as a site of symptom onset-specific checklist of minimal clinical examination, neuroimaging, electrodiagnostic, pulmonary function and laboratory test information required to confirm the diagnosis of ALS ('Versailles minimal dataset'). Although introductory discussions addressed the advantages and disadvantages of earlier diagnosis, false-positive or false-negative diagnosis, the frequency of follow-up and what potential biological markers to be followed, these issues will have to be further evaluated at future consensus conferences.

Diagnosis, Differential↗

A cross-comparison of a large dataset of genes.

SUMMARY: We make available a large cross-comparison for 16 of the completely sequenced genomes and additional eukaryotic genes. The alignments were performed at the protein level using liberal similarity bounds in order to capture as many significant alignments as possible. This dataset will be updated as new genomes become available.

Animals↗

BTS: a scalable Bayesian Tissue Score for prioritizing GWAS variants and their functional contexts across >1000s of omics datasets.

MOTIVATION: statistics from genome-wide association studies (GWAS) are widely used in fine-mapping and colocalization analyses to identify causal variants and their enrichment in functional contexts, such as affected cell types and genomic features. With the expansion of functional genomic (FG) datasets, which now include hundreds of thousands of tracks across various cell and tissue types, it is critical to establish scalable algorithms integrating thousands of diverse FG annotations with GWAS results. RESULTS: We propose BTS (Bayesian Tissue Score), a novel, highly efficient algorithm uniquely designed for (i) identifying affected cell types and functional elements (context-mapping) and (ii) fine-mapping potentially causal variants in a context-specific manner using large collections of cell type-specific FG annotation tracks. BTS leverages GWAS summary statistics and annotation-specific Bayesian models to analyze genome-wide annotation tracks, including enhancers, open chromatin, and histone marks. We evaluated BTS on GWAS summary statistics for immune and cardiovascular traits, such as Inflammatory Bowel Disease (IBD), Rheumatoid Arthritis (RA), Systemic Lupus Erythematosus (SLE), and Coronary Artery Disease (CAD). Our results demonstrate that BTS is over 100× more efficient in estimating functional annotation effects and context-specific variant fine-mapping compared to existing methods. Importantly, this large-scale Bayesian approach prioritizes both known and novel annotations, cell types, genomic regions, and variants and provides valuable biological insights into the functional contexts of these diseases. AVAILABILITY AND IMPLEMENTATION: Docker image is available at https://hub.docker.com/r/wanglab/bts with preinstalled BTS R package (https://bitbucket.org/wanglab-upenn/BTS-R) and BTS GWAS summary statistics analysis pipeline (https://bitbucket.org/wanglab-upenn/bts-pipeline).

Genome-Wide Association Study↗

polars-bio-fast, scalable, and out-of-core operations on large genomic interval datasets.

MOTIVATION: Genomic studies very often rely on computationally intensive analyses of relationships between features, which are typically represented as intervals along a 1D coordinate system (such as positions on a chromosome). In this context, the Python programming language is extensively used for manipulating and analyzing data stored in a tabular form of rows and columns, called a DataFrame. Pandas is the most widely used Python DataFrame package and has been criticized for inefficiencies and scalability issues, which its modern alternative-Polars-aims to address with a native backend written in the Rust programming language. RESULTS: polars-bio is a Python library that enables fast, parallel and out-of-core operations on large genomic interval datasets. Its main components are implemented in Rust, using the Apache DataFusion query engine and Apache Arrow for efficient data representation. It is compatible with Polars and Pandas DataFrame formats. In a real-world comparison (107 versus 1.2×106 intervals), our library runs overlap queries 6.5×, nearest queries 15.5×, count_overlaps queries 38×, and coverage queries 15× faster than Bioframe. On equally sized synthetic sets (107 versus 107), the corresponding speedups are 1.6×, 5.5×, 6×, and 6×. In streaming mode, on real and synthetic interval pairs, our implementation uses 90× and 15× less memory for overlap, 4.5× and 6.5× less for nearest, 60× and 12× less for count_overlaps, and 34× and 7× less for coverage than Bioframe. Multi-threaded benchmarks show good scalability characteristics. To the best of our knowledge, polars-bio is the most efficient single-node library for genomic interval DataFrames in Python. AVAILABILITY AND IMPLEMENTATION: polars-bio is an open-source Python package distributed under the Apache License available for major platforms, including Linux, macOS, and Windows in the PyPI registry. The online documentation is https://biodatageeks.org/polars-bio/ and the source code is available on GitHub: https://github.com/biodatageeks/polars-bio and Zenodo: https://doi.org/10.5281/zenodo.16374290. are available at Bioinformatics online.

Software↗

SpatialRNA: a Python package for easy application of Graph Neural Network models on single-molecule spatial transcriptomics dataset.

SUMMARY: Image-based spatial transcriptomics (iST) deliver gene expression measurements of RNA transcripts in tissue slices with single-molecule resolution and spatial context preserved. Modern Graph Neural Network (GNN) models are promising methods for capturing the complex molecular and cellular phenotypes in tissues at single-transcript and single-cell levels. A key application of GNNs is the detection of spatial domains or niches, that is, groups of molecules and/or cells that collaboratively work together to produce complex phenotypes. Due to the vast number of detected transcripts in (iST) dataset, applying GNNs on RNA molecule graphs is not trivial. We present a Python package, SpatialRNA, for easy (sub)graph generation from tissue samples and provide comprehensive tutorials for convenient and efficient application of Graph Neural Network models under the PyG framework. This highly scalable tool comprehensively segments tissue into spatial domains, aiding in biological interpretation of iST data and its underlying molecular microenvironments. AVAILABILITY AND IMPLEMENTATION: The SpatialRNA package is freely accessible from online repository https://github.com/ruqianl/spatialrna and can be installed via pip. Comprehensive tutorials, guidance on parameter selection, and complete workflows of case studies are available from the documentation website https://ruqianl.github.io/spatialrna_docs/, and uploaded on Zenodo with a DOI 10.5281/zenodo.17339575.

Neural Networks, Computer↗