Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Effect of training datasets on support vector machine prediction of protein-protein interactions.

Knowledge of protein-protein interaction is useful for elucidating protein function via the concept of 'guilt-by-association'. A statistical learning method, Support Vector Machine (SVM), has recently been explored for the prediction of protein-protein interactions using artificial shuffled sequences as hypothetical noninteracting proteins and it has shown promising results (Bock, J. R., Gough, D. A., Bioinformatics 2001, 17, 455-460). It remains unclear however, how the prediction accuracy is affected if real protein sequences are used to represent noninteracting proteins. In this work, this effect is assessed by comparison of the results derived from the use of real protein sequences with that derived from the use of shuffled sequences. The real protein sequences of hypothetical noninteracting proteins are generated from an exclusion analysis in combination with subcellular localization information of interacting proteins found in the Database of Interacting Proteins. Prediction accuracy using real protein sequences is 76.9% compared to 94.1% using artificial shuffled sequences. The discrepancy likely arises from the expected higher level of difficulty for separating two sets of real protein sequences than that for separating a set of real protein sequences from a set of artificial sequences. The use of real protein sequences for training a SVM classification system is expected to give better prediction results in practical cases. This is tested by using both SVM systems for predicting putative protein partners of a set of thioredoxin related proteins. The prediction results are consistent with observations, suggesting that real sequence is more practically useful in development of SVM classification system for facilitating protein-protein interaction prediction.

Algorithms↗

Proteomic dataset of Sca-1+ progenitor cells.

Embryonic stem cells (ES cells) can differentiate into endothelial cells and smooth muscle cells (SMCs), which participate in vascular angiogenesis. In this study, we differentiated mouse ES cells into Sca-1(+) cells, which have the potential to serve as vascular progenitor cells, and mapped their proteome by 2-DE using a pH 3-10 non-linear gradient and 12% SDS-polyacrylamide gels. A subset of 300 protein spots was analysed and mapped, with 241 protein spots being identified by their PMF using MALDI-TOF MS or by partial amino acid sequencing using MS/MS. Our protein map is the first of Sca-1(+) progenitor cells and will facilitate the identification of proteins differentially expressed during stem cell differentiation. The proteome of adult arterial SMCs is described in an accompanying paper (in this issue, DOI 10.1002/pmic.200402045). All data are made accessible on our website http://www.vascular-proteomics.com.

Animals↗

Proteomic dataset of mouse aortic smooth muscle cells.

In an accompanying study (in this issue, DOI 10.1002/pmic.200402044), we have characterised the proteome of Sca-1(+) progenitor cells, which may function as precursors of vascular smooth muscle cells (SMCs). In the present study, we have analysed and mapped protein expression in aortic SMCs of mice, using 2-DE, MALDI-TOF MS and MS/MS. The 2-D system comprised a non-linear immobilised pH 3-10 gradient in the first dimension (separating proteins with pI values of pH 3-10), and 12%T SDS-PAGE in the second dimension (separating proteins in the range 15,000-150,000 Da). Of the 2400 spots visualised, a subset of 267 protein spots was analysed, with 235 protein spots being identified corresponding to 154 unique proteins. The data presented here are the first map of aortic SMCs and the most extensive analysis of SMC proteins published so far. This valuable tool should provide a basis for comparative studies of protein expression in vascular smooth muscle of transgenic mice and is available on our website hhtp://www.vascular-proteomics.com.

Animals↗

A combined dataset of human cerebrospinal fluid proteins identified by multi-dimensional chromatography and tandem mass spectrometry.

Human cerebrospinal fluid (CSF) is an important source for studying protein biomarkers of age-related neurodegenerative diseases. Before characterizing biomarkers unique to each disease, it is necessary to categorize CSF proteins systematically and extensively. However, the enormous complexity, great dynamic range of protein concentrations, and tremendous protein heterogeneity due to post-translational modification of CSF create significant challenges to the existing proteomics technologies for an in-depth, nonbiased profiling of the human CSF proteome. To circumvent these difficulties, in the last few years, we have utilized several different separation methodologies and mass spectrometric platforms that greatly enhanced the identification coverage and the depth of protein profiling of CSF to characterize CSF proteome. In total, 2594 proteins were identified in well-characterized pooled human CSF samples using stringent proteomics criteria. This report summarizes our efforts to comprehensively characterize the human CSF proteome to date.

Cerebrospinal Fluid Proteins↗

Comparison of filtering methods for fMRI datasets.

When studying complex cognitive tasks using functional magnetic resonance imaging (fMRI) one often encounters weak signal responses. These weak responses are corrupted by noise and artifacts of various sources. Preprocessing of the raw data before the application of test statistics helps to extract the signal and can vastly improve signal detection. Artifact sources and algorithms to handle them are discussed. In an empirical approach targeted to yield an optimal recovery of the hemodynamic response, we implemented a test bed for baseline correction and noise-filtering methods. A known signal is modulated onto foreground patches obtained from event-related fMRI experiments. Quantitative performance measures are defined to optimize the characteristics of a given filter and to compare their results. Marked improvements in the sensitivity and selectivity are achieved by optimized filtering. Examples using real data underline the usefulness of this preprocessing sequence.

Algorithms↗

A new approach for improving diagnostic accuracy in Alzheimer's disease and frontal lobe dementia utilising the intrinsic properties of the SPET dataset.

Alzheimer's disease (AD) and frontal lobe dementia (FLD) show characteristic patterns of regional cerebral blood flow (rCBF). However, these patterns may overlap with those observed in the aging brain in elderly normal individuals. The aim of this study was to develop a new method for better classification and recognition of AD and FLD cases as compared with normal controls. Forty-six patients with AD, 7 patients with FLD and 34 normal controls (CTR) were included in the study. rCBF was assessed by technetium-99m hexamethylpropylene amine oxime and a three-headed single-photon emission tomography (SPET) camera. A brain atlas was used to define volumes of interest (VOIs) corresponding to the brain lobes. In addition to conventional image processing methods, based on count density/voxel, the new approach also analysed other intrinsic properties of the data by means of gradient computation steps. Hereby, five factors were assessed and tested separately: the mean count density/voxel and its histogram, the mean gradient and its histogram, and the gradient angle co-occurrence matrix. A feature vector concatenating single features was also created and tested. Preliminary feature discrimination was performed using a two-sided t-test and a K-means clustering was then used to classify the image sets into categories. Finally, five-dimensional co-occurrence matrices combining the different intrinsic properties were computed for each VOI, and their ability to recognise the group to which each individual scan belonged was investigated. For correct classification of the AD-CTR groups, the gradient histogram in the parieto-temporal lobes was the most useful single feature (accuracy 91%). FLD and CTR were better classified by the count density/voxel histogram (frontal and occipital lobes) and by the mean gradient (frontal, temporal and parietal lobes, accuracy 98%). For AD and FLD the count density/voxel histogram in the frontal, parietal and occipital lobes classified the groups with the highest accuracy (85%). The concatenated joint feature correctly classified 96% of the AD-CTR, 98% of the FLD-CTR and 94% of the AD-FLD cases. 5D co-occurrence matrices correctly recognised 98% of the AD-CTR cases, 100% of the FLD-CTR cases and 98% of the AD-FLD cases. The proposed approach classified and diagnosed AD and FLD patients with higher accuracy than conventional analytical methods used for rCBF-SPET. This was achieved by extracting from the SPET data the intrinsic information content in each of the selected VOIs.

Aged↗

In-vivo flow simulation in coronary arteries based on computed tomography datasets: feasibility and initial results.

The purpose of this paper was to non-invasively assess hemodynamic parameters such as mass flow, wall shear stress (WSS), and wall pressure with computational fluid dynamics (CFD) in coronary arteries using patient-specific data from computed tomography (CT) angiography. Five patients (two without atherosclerosis, three with atherosclerosis) underwent retrospectively electrocardiogram (ECG) gated 16-detector row CT using ECG-pulsing and geometric models of coronary arteries were reconstructed for CFD analysis. Blood flow was considered laminar, incompressible, Newtonian, and pulsatile. The mass flow, WSS, and wall pressure were quantified and flow patterns were visualized. The wall pressure continuously decreased towards distal segments and showed pressure drops in stenotic segments. In coronary segments without atherosclerotic wall changes, WSS remained low, even during phases of high flow velocity, whereas in atherosclerotic vessels, the WSS was elevated already at low flow velocities. Stenoses and post-stenotic dilatations led to flow acceleration and rapid deceleration, respectively, including a distortion of flow. Areas of high WSS and high flow velocities were found adjacent to plaques, with values correlating with the degree of stenosis. CFD provided detailed mass flow measurements. CFD analysis is feasible in normal and atherosclerotic coronary arteries and provides the rationale for further investigation of the links between hemodynamic parameters and the significance of coronary stenoses.

Aged↗

Climatic controls of vegetation vigor in four contrasting forest types of India--evaluation from National Oceanic and Atmospheric Administration's Advanced Very High Resolution Radiometer datasets (1990-2000).

Ten-day advanced very high resolution radiometer images from 1990 to 2000 were used to examine spatial patterns in the normalized difference vegetation index (NDVI) and their relationships with climatic variables for four contrasting forest types in India. The NDVI signal has been extracted from homogeneous vegetation patches and has been found to be distinct for deciduous and evergreen forest types, although the mixed-deciduous signal was close to the deciduous ones. To examine the decadal response of the satellite-measured vegetation phenology to climate variability, seven different NDVI metrics were calculated using the 11-year NDVI data. Results suggested strong spatial variability in forest NDVI metrics. Among the forest types studied, wet evergreen forests of north-east India had highest mean NDVI (0.692) followed by evergreen forests of the Western Ghats (0.529), mixed deciduous forests (0.519) and finally dry deciduous forests (0.421). The sum of NDVI (SNDVI) and the time-integrated NDVI followed a similar pattern, although the values for mixed deciduous forests were closer to those for evergreen forests of the Western Ghats. Dry deciduous forests had higher values of inter-annual range (RNDVI) and low mean NDVI, also coinciding with a high SD and thus a high coefficient of variation (CV) in NDVI (CVNDVI). SNDVI has been found to be high for wet evergreen forests of north-east India, followed by evergreen forests of the Western Ghats, mixed deciduous forests and dry deciduous forests. Further, the maximum NDVI values of wet evergreen forests of north-east India (0.624) coincided with relatively high annual total precipitation (2,238.9 mm). The time lags had a strong influence in the correlation coefficients between annual total rainfall and NDVI. The correlation coefficients were found to be comparatively high (R2=0.635) for dry deciduous forests than for evergreen forests and mixed deciduous forests, when the precipitation data with a lag of 30 days was correlated against NDVI. Using multiple regression approach models were developed for individual forest types using 16 different climatic indices. A high proportion of the temporal variance (>90%) has been accounted for by three of the precipitation parameters (maximum precipitation, precipitation of the wettest quarter and driest quarter) and two of the temperature parameters (annual mean temperature and temperature of the coldest quarter) for mixed deciduous forests. Similarly, in the case of deciduous forests, four precipitation parameters and three temperature parameters explained nearly 83.6% of the variance. These results suggest differences in the relationship between NDVI and climatic variables based upon the time of growing season, time interval and climatic indices over which they were summed. These results have implications for forest cover mapping and monitoring in tropical regions of India.

Climate↗

Landscape as a predictor of wetland condition: an evaluation of the Landscape Development Index (LDI) with a large reference wetland dataset from Ohio.

Recent approaches to wetland assessment have advocated a multilevel approach which incorporates assessments based on landscape (remote sensing) data, on-site but "rapid" methods, and intensive methods where quantitative data is collected. Brown and Vivas (2004) recently pro- posed an assessment method that uses remote sensing information (Landscape Development Index or LDI) and propose that it may also be usable as a quantified human disturbance gradient. The LDI was evaluated using a large reference wetland data set from Ohio using land use percentages within a 1 km radius circle of the wetlands. The LDI had interpretable and significant relationships with another human disturbance gradient (the Ohio Rapid Assessment Method for Wetlands or ORAM) and with most metrics and scores from the Vegetation Index of Biotic Integrity (VIBI) developed for use in the State of Ohio. Metrics from emergent wetlands had the most significant correlations with the LDI (10 of 10 metrics), followed by forested wetlands (8 of 10 metrics) and shrub wetlands (4 of 10). Poor correlation for VIBI scores and metrics of shrub wetlands was due to differences in attainable LDI scores based on ecoregion and natural buffers shielding the wetland from otherwise intensive land uses. The ORAM and VIBI were developed for use in wetlands in Ohio completely independent of the LDI. It is an important test of the LDI concept that so many interpretable and significant relationships occurred between the VIBI scores, VIBI metric values, and the ORAM scores. For the purposes of VIBI development, the LDI is an independent, quantified disturbance gradient that has provided an additional test of the VIBI. Given its theoretical underpinnings and the fact that it uses quantified land use percentages, the LDI has many advantages over more qualitative human disturbance gradients. Using land use percentages from increasingly smaller distances from the wetland edge (100-200 m) may improve the resolution of the LDI to detect on-site disturbances to a wetland which degrade its ecological condition. The LDI should be evaluated with other large reference data sets in other regions to evaluate its validity and usefulness as an assessment tool.

Conservation of Natural Resources↗

Using of high-resolution topsoil magnetic screening for assessment of dust deposition: comparison of forest and arable soil datasets.

Magnetic susceptibility (kappa) is an easily detectable geophysical parameter that can be used as a proxy or semi-quantitative tracer of atmospheric industrial and urban dusts deposited in topsoil. An enhanced kappa value of topsoil is in many cases also associated with high concentrations of soil pollutants (mostly heavy metals). High-resolution magnetic screening of topsoil in areas of high pollution influx is a useful tool for detection of pollution "hot spots". General and regional screening maps with a grid density of 10 or 5 km have been performed on the basis of forest topsoil measurement only. The purpose of this study was to perform high-resolution magnetic screening with different grid densities in both forested and agricultural areas (arable land). Our large study area (ca. 200 km(2)) was located in a relatively more polluted region of the central part of Upper Silesia, and a second (small) one (ca. 100 m(2)) was located in the western part of Upper Silesia, with considerably lower influx of pollution. In the framework of this study, we applied a statistical comparison of data obtained in forested areas and on arable land. The arable soil showed statistically significantly lower kappa values, the result of "physical dilution" of the arable layer caused by annual ploughing. Thus arable soils must be avoided during high-resolution field measurement. From semivariograms, it was clear that the spatial correlations in forest topsoil are much stronger than in arable soil, which suggests that a denser measurement grid is required in forested areas.

Dust↗

Induction of comprehensible models for gene expression datasets by subgroup discovery methodology.

Finding disease markers (classifiers) from gene expression data by machine learning algorithms is characterized by a high risk of overfitting the data due the abundance of attributes (simultaneously measured gene expression values) and shortage of available examples (observations). To avoid this pitfall and achieve predictor robustness, state-of-the-art approaches construct complex classifiers that combine relatively weak contributions of up to thousands of genes (attributes) to classify a disease. The complexity of such classifiers limits their transparency and consequently the biological insights they can provide. The goal of this study is to apply to this domain the methodology of constructing simple yet robust logic-based classifiers amenable to direct expert interpretation. On two well-known, publicly available gene expression classification problems, the paper shows the feasibility of this approach, employing a recently developed subgroup discovery methodology. Some of the discovered classifiers allow for novel biological interpretations.

Algorithms↗

Effective detection of remote homologues by searching in sequence dataset of a protein domain fold.

Profile matching methods are commonly used in searches in protein sequence databases to detect evolutionary relationships. We describe here a sensitive protocol, which detects remote similarities by searching in a specialized database of sequences belonging to a fold. We have assessed this protocol by exploring the relationships we detect among sequences known to belong to specific folds. We find that searches within sequences adopting a fold are more effective in detecting remote similarities and evolutionary connections than searches in a database of all sequences. We also discuss the implications of using this strategy to link sequence and structure space.

Databases, Protein↗