Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Towards a consensus on datasets and evaluation metrics for developing B-cell epitope prediction tools.

A B-cell epitope is the three-dimensional structure within an antigen that can be bound to the variable region of an antibody. The prediction of B-cell epitopes is highly desirable for various immunological applications, but has presented a set of unique challenges to the bioinformatics and immunology communities. Improving the accuracy of B-cell epitope prediction methods depends on a community consensus on the data and metrics utilized to develop and evaluate such tools. A workshop, sponsored by the National Institute of Allergy and Infectious Disease (NIAID), was recently held in Washington, DC to discuss the current state of the B-cell epitope prediction field. Many of the currently available tools were surveyed and a set of recommendations was devised to facilitate improvements in the currently existing tools and to expedite future tool development. An underlying theme of the recommendations put forth by the panel is increased collaboration among research groups. By developing common datasets, standardized data formats, and the means with which to consolidate information, we hope to greatly enhance the development of B-cell epitope prediction tools.

Animals↗

Old fold in a new X-ray diffraction dataset? Low-resolution molecular replacement using representative structural templates can provide phase information.

The advent of structural genomics has led to a dramatic increase in the number of structures deposited in the Protein Data Bank. The number of new folds, however, still remains a very small fraction of the total number of deposited structures. Recent data on the progress of the structural genomics initiative reveals that more than 85% of target proteins that progress to the stage of data collection and structure determination have a known fold. Enzymes, which tend to exploit reaction space while adopting a common stable scaffold, contribute significantly to this observation. Herein, we evaluate a method to examine the "old fold in a new dataset" scenario likely to be encountered in the structural genomics pipeline. We demonstrate that a fold detection strategy based on secondary structure signatures followed by molecular replacement using a minimalist model can be effectively used to solve the phase problem in X-ray crystallography without further recourse to heavy atom derivatives or multiple anomalous dispersion techniques. Three common folds-the triosephosphate isomerase (TIM), adenine nucleotide alpha hydrolase-like (HUP), and RNA recognition motif (RRM)-were examined using this approach. The results presented herein also provide an estimate of the extent of phase information that can be derived from a single domain in a large multidomain structure.

Crystallography, X-Ray↗

New method for the analysis of multiple positron emission tomography dynamic datasets: an example applied to the estimation of the cerebral metabolic rate of oxygen.

Positron emission tomography (PET) provides the ability to extract useful quantitative information not available through other radiological techniques. In certain studies, the physiological parameters of interest cannot be determined from the data obtained from a single PET experiment alone. In this case, multiple experiments are required. At present, the methods used to analyse measurements acquired from multiple experiments often involve considering them separately during the modelling procedures. These methods of analysis may cause errors to be propagated through successive modelling procedures and do not fully utilise the information content provided by the PET measurements. A new method is presented, based on linear least squares for the analysis of PET dynamic data acquired from multiple experiments. This method simultaneously considers the complete set of measurements obtained and provides reliable parameter estimates. The efficient use of the information content provided by multiple experiments is considered and the propagation of errors is discussed. To facilitate our discussion, we apply this new method to the estimation of the cerebral metabolic rate of oxygen and the parameters of the oxygen utilisation model as a practical example. The results demonstrate a significant improvement in the reliability and estimation accuracy of the estimates for this new method. Furthermore, this method reduced the likelihood of errors being propagated. Therefore, the proposed method is suitable for the analysis of multiple PET dynamic datasets.

Brain↗

3D visualisation of the middle ear and adjacent structures using reconstructed multi-slice CT datasets, correlating 3D images and virtual endoscopy to the 2D cross-sectional images.

The 3D imaging of the middle ear facilitates better understanding of the patient's anatomy. Cross-sectional slices, however, often allow a more accurate evaluation of anatomical structures, as some detail may be lost through post-processing. In order to demonstrate the advantages of combining both approaches, we performed computed tomography (CT) imaging in two normal and 15 different pathological cases, and the 3D models were correlated to the cross-sectional CT slices. Reconstructed CT datasets were acquired by multi-slice CT. Post-processing was performed using the in-house software "3D Slicer", applying thresholding and manual segmentation. 3D models of the individual anatomical structures were generated and displayed in different colours. The display of relevant anatomical and pathological structures was evaluated in the greyscale 2D slices, 3D images, and the 2D slices showing the segmented 2D anatomy in different colours for each structure. Correlating 2D slices to the 3D models and virtual endoscopy helps to combine the advantages of each method. As generating 3D models can be extremely time-consuming, this approach can be a clinically applicable way of gaining a 3D understanding of the patient's anatomy by using models as a reference. Furthermore, it can help radiologists and otolaryngologists evaluating the 2D slices by adding the correct 3D information that would otherwise have to be mentally integrated. The method can be applied to radiological diagnosis, surgical planning, and especially, to teaching.

Adolescent↗

Analysis of datasets showing which compounds kill which organisms: inferring two systems.

Experiments are sometimes conducted to show which of several compounds is successful at inactivating which of several microorganisms. This paper proposes a method of analysing the datasets obtained. The example refers to 33 compounds (2,4-dihydroxythiobenzanilides) and seven microorganisms (dermatophytes) [data from Eur. J. Med. Chem. 35 (2000) 393]. The conclusion is that two systems or modes of action are needed to explain the data.

Anilides↗

A complete small molecule dataset from the protein data bank.

A complete set of 6300 small molecule ligands was extracted from the protein data bank, and deposited online in PubChem as data source 'SMID'. This set's major improvement over prior methods is the inclusion of cyclic polypeptides and branched polysaccharides, including an unambiguous nomenclature, in addition to normal monomeric ligands. Only the best available example of each ligand structure is retained, and an additional dataset is maintained containing co-ordinates for all examples of each structure. Attempts are made to correct ambiguous atomic elements and other common errors, and a perception algorithm was used to determine bond order and aromaticity when no other information was available.

Databases, Protein↗

Effects of information and machine learning algorithms on word sense disambiguation with small datasets.

Current approaches to word sense disambiguation use (and often combine) various machine learning techniques. Most refer to characteristics of the ambiguity and its surrounding words and are based on thousands of examples. Unfortunately, developing large training sets is burdensome, and in response to this challenge, we investigate the use of symbolic knowledge for small datasets. A naïve Bayes classifier was trained for 15 words with 100 examples for each. Unified Medical Language System (UMLS) semantic types assigned to concepts found in the sentence and relationships between these semantic types form the knowledge base. The most frequent sense of a word served as the baseline. The effect of increasingly accurate symbolic knowledge was evaluated in nine experimental conditions. Performance was measured by accuracy based on 10-fold cross-validation. The best condition used only the semantic types of the words in the sentence. Accuracy was then on average 10% higher than the baseline; however, it varied from 8% deterioration to 29% improvement. To investigate this large variance, we performed several follow-up evaluations, testing additional algorithms (decision tree and neural network), and gold standards (per expert), but the results did not significantly differ. However, we noted a trend that the best disambiguation was found for words that were the least troublesome to the human evaluators. We conclude that neither algorithm nor individual human behavior cause these large differences, but that the structure of the UMLS Metathesaurus (used to represent senses of ambiguous words) contributes to inaccuracies in the gold standard, leading to varied performance of word sense disambiguation techniques.

Algorithms↗

Characterisation of the two most abundant genes in the Haemonchus contortus expressed sequence tag dataset.

Analysis of the Haemonchus contortus Expressed sequence tag (EST) dataset revealed that almost 10% of all ESTs (1719 ESTs) belong to a family of related genes. Close analysis of the ESTs suggests that these represent two genes (called here Hc-nim-1 and Hc-nim-2) with multiple alleles of each. These genes show significant similarity to two genes from Caenorhabditis elegans, F54D5.3 (Wormbase accession WBGene00010049, corresponding protein WP:CE28033) and F54D5.4 (WBGene00010050, WP:CE03409) of unknown function. Reverse transcriptase coupled-PCR showed that both genes are transcribed from the L4 stage onwards and are transcribed in both male and female adult worms. A partial bacterial recombinant of the Hc-NIM-1 protein was made and used to raise antiserum in rabbits which recognised a 19 kDa antigen in the water soluble protein fraction of adult worms. By immunohistochemistry, the Hc-NIM-1 protein was localised in the hypodermis of the pharyngeal region of adult worms but not posterior in the hypodermis surrounding the reproductive tract. To investigate the function of this novel protein family we conducted a RNA interference experiment for the homologuous proteins in C. elegans. No visible phenotype was detected after simultaneous RNAi treatment for both Ce-F54D5.3 and Ce-F54D5.4.

Amino Acid Sequence↗

Benefit of respiration-gated stereotactic radiotherapy for stage I lung cancer: an analysis of 4DCT datasets.

PURPOSE: High local control rates have been reported with stereotactic radiotherapy (SRT) for Stage I non-small-cell lung cancer. Because high-dose fractions are used, reduction in treatment portals will reduce the risk of toxicity to adjacent structures. Respiratory gating can allow reduced field sizes and planning four-dimensional computed tomography scans were retrospectively analyzed to study the benefits for gated SRT and identify patients who derive significant benefit from this approach. METHODS AND MATERIALS: A total of 31 consecutive patients underwent a four-dimensional computed tomography scan, in which three-dimensional computed tomography datasets for 10 phase bins of the respiratory cycle were acquired during free breathing. For a total of 34 tumors, the three planning target volumes (PTVs) were analyzed, namely (1) PTV(10bins), derived from an internal target volume (ITV) that incorporated all observed mobility (ITV(10bins)), with the addition of a 3-mm isotropic setup margin; (2) PTV(gating), derived from an ITV generated from mobility observed in three consecutive phases ("bins") during tidal-expiration, plus addition of a 3-mm isotropic margin; and (3) PTV(10 mm), derived from the addition of a 10-mm isotropic margin to the most central gross tumor volumes in the three bins selected for gating. RESULTS: The PTV(10bins) and PTV(gating) were, on average, 48.2% and 33.3% of the PTV(10 mm), and respective mean volumes of normal tissue (outside the PTV) receiving the prescribed doses were 57.1% and 39.1%, respectively, of that of PTV(10 mm). A significant correlation was seen between the extent of tumor mobility (i.e., a three-dimensional mobility vector of at least 1 cm) and reduction in normal tissue irradiation achieved with gating. The ratio of the intersecting and the encompassing volumes of GTVs at extreme phases of tidal respiration predicted for the benefits of gated respiration. CONCLUSION: The use of "standard population-based" margins for SRT leads to unnecessary normal tissue irradiation. The risk of toxicity is further reduced if respiration-gated radiotherapy is used to treat mobile tumors. These findings suggest that gated SRT will be of clinical relevance in selected patients with mobile tumors.

Carcinoma, Non-Small-Cell Lung↗

Hierarchical modeling of tumor subtypes in cell lines using large-scale genomic datasets.

Cancer cell lines (CLs) are widely used to study tumor biology and drug response, yet their translational relevance is often limited by inaccurate subtype annotations. Existing CL-tumor matching approaches are frequently constrained by flat classification schemes, weak subtype definitions, and the exclusion of normal tissue references, leading to potential confounding of tumor-specific and tissue-of-origin signals. To address these limitations, a hierarchical classification (HC) framework is presented in which CLs are aligned with patient tumors across biological resolutions, from organ to molecular subtype. Gene expression profiles from 802 CLs, 5,612 tumors from The Cancer Genome Atlas (TCGA) , and 8,939 non-cancerous tissues were integrated to separate oncogenic signals from tissue-specific signals. Node-specific features were selected using maximum relevance minimum redundancy, and balanced accuracies of 89% in cross-validation and 75%, and 80% on external datasets were achieved. Through the framework, 43 CLs were reassigned, and clinically relevant underrepresented subtypes were identified.

cancer cell lines↗

A new technique (COMSPARI) to facilitate the identification of minor compounds in complex mixtures by GC/MS and LC/MS: tools for the visualization of matched datasets.

In the rapidly growing field of metabolomics, it is common to analyze complex biological samples by chromatography coupled to mass spectrometry. While several techniques are available for the detection of significant peaks in individual samples, it is still difficult to determine small differences between similar samples. Using conventional software, visual inspections of individual chromatograms or individual mass spectra are often of little use because the differences in the composition of small molecules are too small to be recognizable. Thus, we developed a new approach to visualizing mass spectral datasets using a tool that allows one to easily detect these small differences between mass spectra and chromatograms derived from matched samples. Using these tools on extracts from wild-type and methyltransferase knockout strains of the yeast Saccharomyces cerevisiae, we were able to readily identify those mass spectra in our data sets that were different between the wild-type and the knockout extracts and to identify the molecules involved. The software was also successfully applied to a set of LC/MS data from peptide digests that were performed with identical substrates but different enzymes. We have named this visualization tool COMSPARI (COMparision of SPectrAl Retention Information) and are making the software publicly available via Internet at.

Cell Extracts↗

Bayesian networks for knowledge discovery in large datasets: basics for nurse researchers.

The growth of nursing databases necessitates new approaches to data analyses. These databases, which are known to be massive and multidimensional, easily exceed the capabilities of both human cognition and traditional analytical approaches. One innovative approach, knowledge discovery in large databases (KDD), allows investigators to analyze very large data sets more comprehensively in an automatic or a semi-automatic manner. Among KDD techniques, Bayesian networks, a state-of-the art representation of probabilistic knowledge by a graphical diagram, has emerged in recent years as essential for pattern recognition and classification in the healthcare field. Unlike some data mining techniques, Bayesian networks allow investigators to combine domain knowledge with statistical data, enabling nurse researchers to incorporate clinical and theoretical knowledge into the process of knowledge discovery in large datasets. This tailored discussion presents the basic concepts of Bayesian networks and their use as knowledge discovery tools for nurse researchers.

Artificial Intelligence↗

The use of sparse CT datasets for auto-generating accurate FE models of the femur and pelvis.

The finite element (FE) method when coupled with computed tomography (CT) is a powerful tool in orthopaedic biomechanics. However, substantial data is required for patient-specific modelling. Here we present a new method for generating a FE model with a minimum amount of patient data. Our method uses high order cubic Hermite basis functions for mesh generation and least-square fits the mesh to the dataset. We have tested our method on seven patient data sets obtained from CT assisted osteodensitometry of the proximal femur. Using only 12 CT slices we generated smooth and accurate meshes of the proximal femur with a geometric root mean square (RMS) error of less than 1 mm and peak errors less than 8 mm. To model the complex geometry of the pelvis we developed a hybrid method which supplements sparse patient data with data from the visible human data set. We tested this method on three patient data sets, generating FE meshes of the pelvis using only 10 CT slices with an overall RMS error less than 3 mm. Although we have peak errors about 12 mm in these meshes, they occur relatively far from the region of interest (the acetabulum) and will have minimal effects on the performance of the model. Considering that linear meshes usually require about 70-100 pelvic CT slices (in axial mode) to generate FE models, our method has brought a significant data reduction to the automatic mesh generation step. The method, that is fully automated except for a semi-automatic bone/tissue boundary extraction part, will bring the benefits of FE methods to the clinical environment with much reduced radiation risks and data requirement.

Algorithms↗

Comparing statistical and semantic approaches for identifying change from land cover datasets.

In this paper, we examine methods for integrating spatial data which apparently should be comparable because they are of the same data type or theme, but which are incompatible or discordant because the classes of that theme are different. For a variety of reasons including changes in methods, in understanding of the resource, and in policy initiatives in the commissioning of the survey, this problem is widespread in the results of natural resources surveys. We present two generic methods: one method is grounded in a statistical approach using discriminant analysis, and the other exploits the knowledge of experts. We use the context of land cover mapping of Great Britain to explore these approaches for integrating discordant data. We demonstrate that the expert-based approach gives very good levels of identification of locations with incompatible classifications at different times, and gives a much better rate of recognition of change. Some conclusions are made about the need to expand current metadata and data quality reporting to include descriptions of:- data conceptualisations, semantics and ontologies;- who decided and defined what the features of interest in a dataset are, and why. If the benefits of spatial data initiatives such as GRID, E-science and INSPIRE are to be fully realised then some method needs to be found to communicate that information most effectively to the potential user of the data.

Conservation of Natural Resources↗

Ventricular shape visualization using selective volume rendering of cardiac datasets.

In this paper, we present a novel technique of improving volume rendering quality and speed by integrating original volume data and global model information attained by segmentation. The segmentation information prevents object occlusions that may appear when volume rendering is based on local image features only. Thus the presented visualization technique provides meaningful visual results that enable a clear understanding of complex anatomical structures. In the first part, we describe a segmentation technique for extracting the region of interest based on an active contour model. In the second part, we propose a volume rendering method for visualizing the selected portions of fuzzy surfaces extracted by local image processing methods. We show the results of selective volume rendering of left and right ventricle based on cardiac datasets from clinical routines. Our method offers an accelerated technique to accurately visualize the surfaces of segmented objects.

Algorithms↗

Analytical methods for studying the evolution of paralogs using duplicate gene datasets.

Gene duplication is widely viewed as an important source of raw material for functional innovation in proteins because at least some duplicate copies will evolve new or slightly modified functions. The study of the molecular processes by which functional innovation occurs interests both evolutionary biologists and protein chemists, and the development of methods to investigate these processes has led to a productive meeting of disciplines and an availability of complementary approaches for exploring datasets. This has resulted in insights into past events, prediction of current function, and prediction of future change. The methods fall broadly into two categories: those that rely on detection of shifts in selective constraints and those that rely on detection of correlations between molecular changes and functional shifts. Strengths and limitations of the methods are evaluated here in the context of the question being addressed, the input required, and the specific metric that is evaluated in each test.

Databases, Genetic↗

Combining independent component analysis and correlation analysis to probe interregional connectivity in fMRI task activation datasets.

A new approach in studying interregional functional connectivity using functional magnetic resonance imaging (fMRI) is presented. Functional connectivity may be detected by means of cross correlating time course data from functionally related brain regions. These data exhibit high temporal coherence of low frequency fluctuations due to synchronized blood flow changes. In the past, this fMRI technique for studying functional connectivity has been applied to subjects that performed no prescribed task ("resting" state). This paper presents the results of applying the same method to task-related activation datasets. Functional connectivity analysis is first performed in areas not involved with the task. Then a method is devised to remove the effects of activation from the data using independent component analysis (ICA) and functional connectivity analysis is repeated. Functional connectivity, which is demonstrated in the "resting brain," is not affected by tasks which activate unrelated brain regions. In addition, ICA effectively removes activation from the data and may allow us to study functional connectivity even in the activated regions.

Adult↗

Analysis, statistical validation and dissemination of large-scale proteomics datasets generated by tandem MS.

Tandem mass spectrometry has been used increasingly for high-throughput analysis of complex protein samples. A major challenge lies in the consistent, objective and transparent analysis of the large amounts of data generated by such experiments and in their dissemination and publication. Here, we review currently available computational tools and discuss the need for statistical criteria in the analysis of large proteomics datasets.

Amino Acid Sequence↗