Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,081 records · Page 60Linked to original sources

The use of anonymous DNA markers in assessing worldwide relatedness in the yeast species Pichia kluyveri Bedford and Kudrjavzev.

Pichia kluyveri, a sexual ascomycetous yeast from cactus necroses and acidic fruit, is divided into three varieties. We used physiological, RAPD, and AFLP data to compare 46 P. kluyveri strains collected worldwide to investigate relationships among varieties. Physiology did not place all strains into described varieties. Although the combined AFLP and RAPD data produced a single most parsimonious tree, separate analysis of AFLP and RAPD data resulted in significantly different trees (by the partition homogeneity test). We then compared the distribution of strains per band to an expected distribution. This suggested we could separate both the AFLP and RAPD datasets into bands from rapidly and slowly changing DNA regions. When only bands from slowly changing regions (from each dataset) were included in the analysis, both the RAPD and AFLP datasets supported a single tree. This second tree did not differ significantly from the cladogram based on all of the DNA data, which we accepted as the best estimate of the phylogeny of these yeast strains. Based on this phylogeny, we were able to demonstrate the strong influence of geography on the population structure of this yeast, confirm the monophyly of one variety, question the utility of maintaining another variety, and demonstrate that the physiological differences used to separate the varieties did not do so in all cases.

Americas↗

Two-stage multi-class support vector machines to protein secondary structure prediction.

Bioinformatics techniques to protein secondary structure (PSS) prediction are mostly single-stage approaches in the sense that they predict secondary structures of proteins by taking into account only the contextual information in amino acid sequences. In this paper, we propose two-stage Multi-class Support Vector Machine (MSVM) approach where a MSVM predictor is introduced to the output of the first stage MSVM to capture the sequential relationship among secondary structure elements for the prediction. By using position specific scoring matrices, generated by PSI-BLAST, the two-stage MSVM approach achieves Q3 accuracies of 78.0% and 76.3% on the RS126 dataset of 126 nonhomologous globular proteins and the CB396 dataset of 396 nonhomologous proteins, respectively, which are better than the highest scores published on both datasets to date.

Computational Biology↗

Discovery of binding motif pairs from protein complex structural data and protein interaction sequence data.

Unravelling the underlying mechanisms of protein interactions requires knowledge about the interactions' binding sites. In this paper, we use a novel concept, binding motif pairs, to describe binding sites. A binding motif pair consists of two motifs each derived from one side of the binding protein sequences. The discovery is a directed approach that uses a combination of two data sources: 3-D structures of protein complexes and sequences of interacting proteins. We first extract maximal contact segment pairs from the protein complexes' structural data. We then use these segment pairs as templates to sub-group the interacting protein sequence dataset, and conduct an iterative refinement to derive significant binding motif pairs. This combination approach is efficient in handling large datasets of protein interactions. From a dataset of 78,390 protein interactions, we have discovered 896 significant binding motif pairs. The discovered motif pairs include many novel motif pairs as well as motifs that agree well with experimentally validated patterns in the literature.

Amino Acid Motifs↗

Comparison of psychiatric ICD-10 diagnoses in Denmark and Germany.

A dataset of psychiatric ICD-10 diagnoses from the Danish case register concerning psychiatric hospitals was compared with a sample of psychiatric diagnoses from 27 psychiatric hospitals in Germany. The comparison shows a higher proportion of F1 diagnoses in the German dataset and a difference in the coding of alcohol dependence and harmful use. Some further differences in the groups F0-F6 are demonstrated and some of them are discussed. The most frequent diagnoses found in both datasets but in different sequence are alcohol dependence syndrome and paranoid schizophrenia and, in third place, adjustment disorder. Various aspects of the problem of rarely used diagnoses are discussed.

Adjustment Disorders↗

apoB-100 has a pentapartite structure composed of three amphipathic alpha-helical domains alternating with two amphipathic beta-strand domains. Detection by the computer program LOCATE.

Due to the great length of apolipoprotein (apo) B-100, the localization of lipid-associating domains in this protein has been difficult. To address this question, we developed a computer program called Locate that searches amino acid sequences to identify potential amphipathic alpha-helixes and beta-strands by using sets of rules for helix and strand termination. A series of model chimeric protein test datasets were created by tandem linking of amino acid sequences of multiple proteins containing four different secondary structural motifs: motif A (exchangeable plasma apolipoproteins); motif G (globular alpha-helical proteins); motif C (coiled-coil alpha-helical proteins); and motif B (beta pleated-sheet proteins). These four test datasets, as well as randomly scrambled sequences of each dataset, were analyzed by Locate using increasingly stringent parameters. Using intermediately stringent parameters under which significant numbers of amphipathic helixes were found only in the unscrambled motif A, two dense clusters of putative lipid-associating amphipathic helixes were located precisely in the middle and at the C-terminal end of apoB-100 (a sparse cluster of class G* helixes is located at the N-terminus). The dense clusters are located between residues 2103 through 2560 and 4061 through 4338 and have densities of 2.4 and 2.2 amphipathic helixes per 100 residues, respectively; under these conditions, motif A has a density of 1.4 amphipathic helixes per 100 residues. These two domains correspond closely to the two major apoB-100 lipid-associated domains at residues 2100 through 2700 and 4100 through 4500 using the principle of releasability of tryptic peptides from trypsin-treated intact low-density lipoprotein. The classes of amphipathic helixes identified within these two putative lipid-associating domains are considerably more diverse than those found in the exchangeable plasma apolipoproteins. Interestingly, apoB-48 terminates at the N-terminal edge of the middle cluster. By using a similar strategy for analysis of amphipathic beta-strands, we discovered that the two gap regions between the three amphipathic helix clusters are highly enriched in putative amphipathic beta-strands, while the three amphipathic helical domains are essentially devoid of this putative lipid-associating motif. We propose, therefore, that apoB-100 has a pentapartite structure, NH2-alpha 1-beta 1-alpha 2-beta 2-alpha 3-COOH, with alpha 1 representing a globular domain.

Algorithms↗

Automated measurement of microaneurysm turnover.

PURPOSE: An automated system for the measurement of microaneurysm (MA) turnover was developed and compared with manual measurement. The system analyses serial fluorescein angiogram (FA) or red-free (RF) fundus images; fluorescein angiography was used in this study because it is the more sensitive test for MAs. Previous studies have shown that the absolute number of MAs observed does not reflect the dynamic temporal nature of the MA population. In this study, almost half of the MAs present at baseline had regressed after a year and been replaced by new lesions elsewhere. METHODS: Two clinical datasets were used to evaluate the performance of the automated turnover measurement system. The first consisted of 10 patients who had two fluorescein angiograms acquired a year apart. These data were analyzed, both manually and using the automated system, to investigate the inter- and intraobserver variations associated with manual measurement and to assess the performance of the automated system. The second dataset contained FAs from a further 25 patients. This dataset was analyzed only with the automated system to investigate some properties of microaneurysm turnover, in particular the differing detection sensitivities of new, static and regressed microaneurysms. RESULTS: Manual measurements exhibited large inter- and intraobserver variation. The sensitivity and specificity of the automated system were similar to those of the human observers. However, the automated measurements were more consistent-an important condition for accurate turnover quantification. Regressed MAs were more difficult to detect reliably than new MAs, which were themselves more difficult to detect reliably than static MAs. CONCLUSIONS: The automated system was shown to be fast, reliable, and repeatable, making it suitable for processing large numbers of images. Performance was similar to that of trained manual observers.

Adult↗

Use of mark-recapture techniques to estimate the size of hard-to-reach populations.

OBJECTIVES: To review the main problems associated with mark-recapture methods of population estimation, and to indicate some practical strategies for addressing these problems, with illustrations from a study of drug-use prevalence across Wales. METHODS: Unnamed identifier data were collected in 1994 on 2610 drug users who were in contact with various agencies across Wales: the police, drug treatment agencies, needle exchanges, probation services, and agencies reporting to the Welsh Drugs Misuse Database. RESULTS: Based on the dependency relationships between different agencies' datasets, different estimates of the 'hidden' populations (not in contact with agencies) were modelled for each county, for males and females, for injecting drug users and serious drug users, and for those under 25 and those over 25 years of age. Different models were also constructed for the same subpopulations, using different agency datasets and different criteria of overlap between them, yielding a total of 230 different models. CONCLUSIONS: The issues of sample heterogeneity and population definition are particularly intractable in mark-recapture studies. Sample heterogeneity may be partly addressed by separately modelling different subpopulations to check whether they show the same dependency relationships as the main population. Population definition may be partly addressed by restricting modelling to datasets thought to share roughly congruent population definitions.

Adult↗

New post-imaging software provides fast and accurate volume data from CTA surveillance after endovascular aneurysm repair.

PURPOSE: To quantify intra- and interobserver variabilities when measuring total aneurysm volume after endovascular aneurysm repair using the Vitrea 2 System and to compare it in terms of accuracy and processing time with the gold standard methods using the Easy Vision workstation. METHODS: Total aneurysm volumes from 30 postendograft CTA datasets were randomly selected from a database consisting of approximately 400 CTA datasets recorded in 89 patients. The intra- and interobserver variabilities were measured on the Vitrea workstation by 2 investigators. The intermodality variability was calculated for the same measurements using the Easy Vision workstation. The differences of each pair of measurements were plotted against their mean, and the repeatability coefficient (RC) was calculated. The mean differences were also expressed as a percentage of the first measurements. RESULTS: The intraobserver mean difference was 1.6 mL (1.4%) with an RC of 10.8 mL (10.1%) and the interobserver mean difference was -1.4 mL (-1.4%) with an RC of 11.7 mL (10.2%). The intermodality mean difference was 1.8 mL (2.0%) with an RC of 15.8 mL (11.1%). The Vitrea workstation required a median of 8 minutes (interquartile range 7-10) for 1 observer and 6 minutes (interquartile range 5-8) for the other to perform a complete volume segmentation of each patient dataset compared to an estimated average of 30 minutes using the Easy Vision workstation. CONCLUSIONS: The Vitrea workstation provides fast and accurate volume data from spiral CTA follow-up of endovascular aneurysm repair. This software may enhance the acceptability of volume surveillance in daily practice.

Aged↗

Repositioning accuracy of a commercially available double-vacuum whole body immobilization system for stereotactic body radiation therapy.

We evaluated the repositioning accuracy of a commercially available stereotactic whole body immobilization system (BodyFIX, Medical Intelligence, Schwabmuenchen, Germany) in 36 patients treated by hypofractionated stereotactic body radiation therapy. CT data were acquired for positional control of patient and tumor before each fraction of the treatment course. Those control CT datasets were compared with the original treatment planning CT simulation and analyzed with respect to positional misalignment of bony patient anatomy, and the respective position of the treated small lung or liver lesions. We assessed the stereotactic coordinates of distinct bony anatomical landmarks in the original CT and each control dataset. In addition, the target isocenter was recorded in the planning CT simulation dataset. An iterative optimization algorithm was implemented, utilizing a root mean square scoring function to determine the best-fit orientation of subsequent sets of anatomical landmark measurements relative to the original treatment planning CT data set. This allowed for the calculation of the x, y and z-components of translation of the patient's body and the target's center-of-mass for each control CT study, as well as rotation about the principal room axes in the respective CT data sets. In addition to absolute patient/target translation, the total magnitude vector of patient and target misalignment was calculated. A clinical assessment determined whether or not the assigned planning target volume safety margins would have provided the desired target coverage. To this end, each control CT study was co-registered with the original treatment planning study using immobilization system related fiducial markers, and the computed isodose calculation was superimposed. In 109 control setup CT scans available for comparison with their respective treatment planning CT simulation study (2-5 per patient, median 3), anatomical landmark analysis revealed a mean bony landmark translation of -0.4 +/- 3.9 (mean +/- SD), -0.1 +/- 1.6 and 0.3 +/- 3.6 mm in x, y and z-directions, respectively. Bony landmark setup deviations along one or more principal axis larger than 5 mm were observed in 32 control CT studies (29.4%). Body rotations about the x-, y- and z-axis were 0.9 +/- 0.7, 0.8 +/- 0.7 and 1.8 +/- 1.6 degrees, respectively. Assuming a rigid body relationship of target and bony anatomy, the mean computed absolute target translation was 2.9 +/- 3.3, 2.3 +/- 2.5 and 3.2 +/- 2.7 mm in x, y and z-directions, respectively. The median and mean magnitude vector of target isocenter displacement was computed to be 4.9 mm, and 5.7 +/- 3.7 mm. Clinical assessment of PTV/target volume coverage revealed 72 (66.1%), 23 (21.1%), and 14 (12.8%), of excellent (100% isodose coverage), good (>90% isodose coverage), and poor GTV/isodose alignment quality (less than 90% isodose coverage to some aspect of the GTV), respectively. Loss of target volume dose coverage was correlated with translations >5 mm along one or more axes (p<0.0001), rotations >3 degrees about the z-axis (p=0.0007) and body mass index >30 (p<0.0001). The analyzed BodyFIX whole body immobilization system performed favorably compared with other stereotactic body immobilization systems for which peer-reviewed repositioning data exist. While the measured variability in patient and target setup provided clinically acceptable setup accuracy in the vast majority of cases, larger setup deviations were occasional observed. Such deviations constitute a potential for partial target underdosing warranting, in our opinion, a pre-delivery positional assessment procedure (e.g., pre-treatment control CT scan).

Body Weights and Measures↗

Members of the glutathione and ABC-transporter families are associated with clinical outcome in patients with diffuse large B-cell lymphoma.

Standard chemotherapy fails in 40% to 50% of patients with diffuse large B-cell lymphoma (DLBCL). Some of these failures can be salvaged with high-dose regimens, suggesting a role for drug resistance in this disease. We examined the expression of genes in the glutathione (GSH) and ATP-dependent transporter (ABC) families in 2 independent tissue-based expression microarray datasets obtained prior to therapy from patients with DLBCL. Among genes in the GSH family, glutathione peroxidase 1 (GPX1) had the most significant adverse effect on disease-specific overall survival (dOS) in the primary dataset (n = 130) (HR: 1.68; 95% CI: 1.26-2.22; P < .001). This effect remained statistically significant after controlling for biologic signature, LLMPP cell-of-origin signature, and IPI score, and was confirmed in the validation dataset (n = 39) (HR: 1.7; 95% CI: 1.05-2.8; P = .033). Recursive partitioning identified a group of patients with low-level expression of GPX1 and multidrug resistance 1 (MDR1; ABCB1) without early treatment failures and with superior dOS (P < .001). Overall, our findings suggest an important association of oxidative-stress defense and drug elimination with treatment failure in DLBCL and identify GPX1 and ABCB1 as potentially powerful biomarkers of early failure and disease-specific survival.

ATP Binding Cassette Transporter, Subfamily B↗

Significance analysis of lexical bias in microarray data.

BACKGROUND: Genes that are determined to be significantly differentially regulated in microarray analyses often appear to have functional commonalities, such as being components of the same biochemical pathway. This results in certain words being under- or overrepresented in the list of genes. Distinguishing between biologically meaningful trends and artifacts of annotation and analysis procedures is of the utmost importance, as only true biological trends are of interest for further experimentation. A number of sophisticated methods for identification of significant lexical trends are currently available, but these methods are generally too cumbersome for practical use by most microarray users. RESULTS: We have developed a tool, LACK, for calculating the statistical significance of apparent lexical bias in microarray datasets. The frequency of a user-specified list of search terms in a list of genes which are differentially regulated is assessed for statistical significance by comparison to randomly generated datasets. The simplicity of the input files and user interface targets the average microarray user who wishes to have a statistical measure of apparent lexical trends in analyzed datasets without the need for bioinformatics skills. The software is available as Perl source or a Windows executable. CONCLUSION: We have used LACK in our laboratory to generate biological hypotheses based on our microarray data. We demonstrate the program's utility using an example in which we confirm significant upregulation of SPI-2 pathogenicity island of Salmonella enterica serovar Typhimurium by the cation chelator dipyridyl.

2,2'-Dipyridyl↗

Estimating mutual information using B-spline functions--an improved similarity measure for analysing gene expression data.

BACKGROUND: The information theoretic concept of mutual information provides a general framework to evaluate dependencies between variables. In the context of the clustering of genes with similar patterns of expression it has been suggested as a general quantity of similarity to extend commonly used linear measures. Since mutual information is defined in terms of discrete variables, its application to continuous data requires the use of binning procedures, which can lead to significant numerical errors for datasets of small or moderate size. RESULTS: In this work, we propose a method for the numerical estimation of mutual information from continuous data. We investigate the characteristic properties arising from the application of our algorithm and show that our approach outperforms commonly used algorithms: The significance, as a measure of the power of distinction from random correlation, is significantly increased. This concept is subsequently illustrated on two large-scale gene expression datasets and the results are compared to those obtained using other similarity measures.A C++ source code of our algorithm is available for non-commercial use from kloska@scienion.de upon request. CONCLUSION: The utilisation of mutual information as similarity measure enables the detection of non-linear correlations in gene expression datasets. Frequently applied linear correlation measures, which are often used on an ad-hoc basis without further justification, are thereby extended.

Algorithms↗

Performance of a genetic algorithm for mass spectrometry proteomics.

BACKGROUND: Recently, mass spectrometry data have been mined using a genetic algorithm to produce discriminatory models that distinguish healthy individuals from those with cancer. This algorithm is the basis for claims of 100% sensitivity and specificity in two related publicly available datasets. To date, no detailed attempts have been made to explore the properties of this genetic algorithm within proteomic applications. Here the algorithm's performance on these datasets is evaluated relative to other methods. RESULTS: In reproducing the method, some modifications of the algorithm as it is described are necessary to get good performance. After modification, a cross-validation approach to model selection is used. The overall classification accuracy is comparable though not superior to other approaches considered. Also, some aspects of the process rely upon random sampling and thus for a fixed dataset the algorithm can produce many different models. This raises questions about how to choose among competing models. How this choice is made is important for interpreting sensitivity and specificity results as merely choosing the model with lowest test set error rate leads to overestimates of model performance. CONCLUSIONS: The algorithm needs to be modified to reduce variability and care must be taken in how to choose among competing models. Results derived from this algorithm must be accompanied by a full description of model selection procedures to give confidence that the reported accuracy is not overstated.

Algorithms↗

Identifying spatially similar gene expression patterns in early stage fruit fly embryo images: binary feature versus invariant moment digital representations.

BACKGROUND: Modern developmental biology relies heavily on the analysis of embryonic gene expression patterns. Investigators manually inspect hundreds or thousands of expression patterns to identify those that are spatially similar and to ultimately infer potential gene interactions. However, the rapid accumulation of gene expression pattern data over the last two decades, facilitated by high-throughput techniques, has produced a need for the development of efficient approaches for direct comparison of images, rather than their textual descriptions, to identify spatially similar expression patterns. RESULTS: The effectiveness of the Binary Feature Vector (BFV) and Invariant Moment Vector (IMV) based digital representations of the gene expression patterns in finding biologically meaningful patterns was compared for a small (226 images) and a large (1819 images) dataset. For each dataset, an ordered list of images, with respect to a query image, was generated to identify overlapping and similar gene expression patterns, in a manner comparable to what a developmental biologist might do. The results showed that the BFV representation consistently outperforms the IMV representation in finding biologically meaningful matches when spatial overlap of the gene expression pattern and the genes involved are considered. Furthermore, we explored the value of conducting image-content based searches in a dataset where individual expression components (or domains) of multi-domain expression patterns were also included separately. We found that this technique improves performance of both IMV and BFV based searches. CONCLUSIONS: We conclude that the BFV representation consistently produces a more extensive and better list of biologically useful patterns than the IMV representation. The high quality of results obtained scales well as the search database becomes larger, which encourages efforts to build automated image query and retrieval systems for spatial gene expression patterns.

Algorithms↗

Integrative analysis of multiple gene expression profiles with quality-adjusted effect size models.

BACKGROUND: With the explosion of microarray studies, an enormous amount of data is being produced. Systematic integration of gene expression data from different sources increases statistical power of detecting differentially expressed genes and allows assessment of heterogeneity. The challenge, however, is in designing and implementing efficient analytic methodologies for combination of data generated by different research groups. RESULTS: We extended traditional effect size models to combine information from different microarray datasets by incorporating a quality measure for each gene in each study into the effect size estimation. We illustrated our method by integrating two datasets generated using different Affymetrix oligonucleotide types. Our results indicate that the proposed quality-adjusted weighting strategy for modelling inter-study variation of gene expression profiles not only increases consistency and decreases heterogeneous results between these two datasets, but also identifies many more differentially expressed genes than methods proposed previously. CONCLUSION: Data integration and synthesis is becoming increasingly important. We live in a high-throughput era where technologies constantly change leaving behind a trail of data with different forms, shapes and sizes. Statistical and computational methodologies are therefore critical for extracting the most out of these related but not identical sources of data.

Algorithms↗

SpectralNET--an application for spectral graph analysis and visualization.

BACKGROUND: Graph theory provides a computational framework for modeling a variety of datasets including those emerging from genomics, proteomics, and chemical genetics. Networks of genes, proteins, small molecules, or other objects of study can be represented as graphs of nodes (vertices) and interactions (edges) that can carry different weights. SpectralNET is a flexible application for analyzing and visualizing these biological and chemical networks. RESULTS: Available both as a standalone .NET executable and as an ASP.NET web application, SpectralNET was designed specifically with the analysis of graph-theoretic metrics in mind, a computational task not easily accessible using currently available applications. Users can choose either to upload a network for analysis using a variety of input formats, or to have SpectralNET generate an idealized random network for comparison to a real-world dataset. Whichever graph-generation method is used, SpectralNET displays detailed information about each connected component of the graph, including graphs of degree distribution, clustering coefficient by degree, and average distance by degree. In addition, extensive information about the selected vertex is shown, including degree, clustering coefficient, various distance metrics, and the corresponding components of the adjacency, Laplacian, and normalized Laplacian eigenvectors. SpectralNET also displays several graph visualizations, including a linear dimensionality reduction for uploaded datasets (Principal Components Analysis) and a non-linear dimensionality reduction that provides an elegant view of global graph structure (Laplacian eigenvectors). CONCLUSION: SpectralNET provides an easily accessible means of analyzing graph-theoretic metrics for data modeling and dimensionality reduction. SpectralNET is publicly available as both a .NET application and an ASP.NET web application from http://chembank.broad.harvard.edu/resources/. Source code is available upon request.

Algorithms↗

Discover protein sequence signatures from protein-protein interaction data.

BACKGROUND: The development of high-throughput technologies such as yeast two-hybrid systems and mass spectrometry technologies has made it possible to generate large protein-protein interaction (PPI) datasets. Mining these datasets for underlying biological knowledge has, however, remained a challenge. RESULTS: A total of 3108 sequence signatures were found, each of which was shared by a set of guest proteins interacting with one of 944 host proteins in Saccharomyces cerevisiae genome. Approximately 94% of these sequence signatures matched entries in InterPro member databases. We identified 84 distinct sequence signatures from the remaining 172 unknown signatures. The signature sharing information was then applied in predicting sub-cellular localization of yeast proteins and the novel signatures were used in identifying possible interacting sites. CONCLUSION: We reported a method of PPI data mining that facilitated the discovery of novel sequence signatures using a large PPI dataset from S. cerevisiae genome as input. The fact that 94% of discovered signatures were known validated the ability of the approach to identify large numbers of signatures from PPI data. The significance of these discovered signatures was demonstrated by their application in predicting sub-cellular localizations and identifying potential interaction binding sites of yeast proteins.

Binding Sites↗

A high level interface to SCOP and ASTRAL implemented in python.

BACKGROUND: Benchmarking algorithms in structural bioinformatics often involves the construction of datasets of proteins with given sequence and structural properties. The SCOP database is a manually curated structural classification which groups together proteins on the basis of structural similarity. The ASTRAL compendium provides non redundant subsets of SCOP domains on the basis of sequence similarity such that no two domains in a given subset share more than a defined degree of sequence similarity. Taken together these two resources provide a 'ground truth' for assessing structural bioinformatics algorithms. We present a small and easy to use API written in python to enable construction of datasets from these resources. RESULTS: We have designed a set of python modules to provide an abstraction of the SCOP and ASTRAL databases. The modules are designed to work as part of the Biopython distribution. Python users can now manipulate and use the SCOP hierarchy from within python programs, and use ASTRAL to return sequences of domains in SCOP, as well as clustered representations of SCOP from ASTRAL. CONCLUSION: The modules make the analysis and generation of datasets for use in structural genomics easier and more principled.

Database Management Systems↗