Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,243 records · Page 69Linked to original sources

Large-scale plant protein subcellular location prediction.

Current plant genome sequencing projects have called for development of novel and powerful high throughput tools for timely annotating the subcellular location of uncharacterized plant proteins. In view of this, an ensemble classifier, Plant-PLoc, formed by fusing many basic individual classifiers, has been developed for large-scale subcellular location prediction for plant proteins. Each of the basic classifiers was engineered by the K-Nearest Neighbor (KNN) rule. Plant-PLoc discriminates plant proteins among the following 11 subcellular locations: (1) cell wall, (2) chloroplast, (3) cytoplasm, (4) endoplasmic reticulum, (5) extracell, (6) mitochondrion, (7) nucleus, (8) peroxisome, (9) plasma membrane, (10) plastid, and (11) vacuole. As a demonstration, predictions were performed on a stringent benchmark dataset in which none of the proteins included has > or =25% sequence identity to any other in a same subcellular location to avoid the homology bias. The overall success rate thus obtained was 32-51% higher than the rates obtained by the previous methods on the same benchmark dataset. The essence of Plant-PLoc in enhancing the prediction quality and its significance in biological applications are discussed. Plant-PLoc is accessible to public as a free web-server at: (http://202.120.37.186/bioinf/plant). Furthermore, for public convenience, results predicted by Plant-PLoc have been provided in a downloadable file at the same website for all plant protein entries in the Swiss-Prot database that do not have subcellular location annotations, or are annotated as being uncertain. The large-scale results will be updated twice a year to include new entries of plant proteins and reflect the continuous development of Plant-PLoc.

Plant Proteins↗

A Causal Effect of Serum 25(OH)D Level on Appendicular Muscle Mass: Evidence From NHANES Data and Mendelian Randomization Analyses.

BACKGROUND: Low serum vitamin D status was reported to be associated with reduced muscle mass; however, it is inconclusive whether this relationship is causal. This study used data from the National Health and Nutrition Examination Survey (NHANES) and two-sample Mendelian randomization (MR) analyses to ascertain the causal relationship between serum 25-hydroxyvitamin D [25(OH)D] and appendicular muscle mass (AMM). METHODS: In the NHANES 2011-2018 dataset, 11&#x2009;242 participants (5588 males and 5654 females) aged 18-59&#x2009;years old were included, and multivariant linear regression was performed to assess the relationship between 25(OH)D and AMM measured by dual-energy X-ray absorptiometry. In two-sample MR analysis, 167 single nucleotide polymorphisms significantly associated with serum 25(OH)D at the genome-wide association level (p&#x2009;<&#x2009;5&#x2009;&#xd7;&#x2009;10-8) were applied as instrumental variables (IVs) to assess vitamin D effects on AMM in the UK Biobank (417&#x2009;580 Europeans) using univariable and multivariable MR (MVMR) models. RESULTS: In the NHANES dataset, serum 25(OH)D concentrations were positively associated with AMM (&#x3b2;&#x2009;=&#x2009;0.013, SE&#x2009;=&#x2009;0.001, p&#x2009;<&#x2009;0.001) in all participants, after adjustment for age, race, season of blood collection, education, income, body mass index and physical activity. In stratification analysis by sex, males (&#x3b2;&#x2009;=&#x2009;0.024, SE&#x2009;=&#x2009;0.002, p&#x2009;<&#x2009;0.001) showed more pronounced positive associations than females (&#x3b2;&#x2009;=&#x2009;0.003, SE&#x2009;=&#x2009;0.002, p&#x2009;=&#x2009;0.024). In univariable MR, genetically higher serum 25(OH)D levels were positively associated with AMM in all participants (&#x3b2;&#x2009;=&#x2009;0.049, SE&#x2009;=&#x2009;0.024, p&#x2009;=&#x2009;0.039) and males (&#x3b2;&#x2009;=&#x2009;0.057, SE&#x2009;=&#x2009;0.025, p&#x2009;=&#x2009;0.021), but only marginally significant in females (&#x3b2;&#x2009;=&#x2009;0.043, SE&#x2009;=&#x2009;0.025, p&#x2009;=&#x2009;0.090) based on IVW models was noticed. No significant pleiotropy effects were detected for the IVs in the two-sample MR investigations. In MVMR analysis, a positive causal effect of 25(OH)D on AMM was observed in the total population (&#x3b2;&#x2009;=&#x2009;0.116, SE&#x2009;=&#x2009;0.051, p&#x2009;=&#x2009;0.022), males (&#x3b2;&#x2009;=&#x2009;0.111, SE&#x2009;=&#x2009;0.053, p&#x2009;=&#x2009;0.036) and females (&#x3b2;&#x2009;=&#x2009;0.124, SE&#x2009;=&#x2009;0.054, p&#x2009;=&#x2009;0.021). CONCLUSIONS: Our results suggested a positive causal effect of serum 25(OH)D concentration on AMM; however, more researches are warranted to unveil the underlying biological mechanisms and evaluate the effects of vitamin D intervention on AMM.

Humans↗

Inferring species phylogenies from multiple genes: concatenated sequence tree versus consensus gene tree.

Phylogenetic trees from multiple genes can be obtained in two fundamentally different ways. In one, gene sequences are concatenated into a super-gene alignment, which is then analyzed to generate the species tree. In the other, phylogenies are inferred separately from each gene, and a consensus of these gene phylogenies is used to represent the species tree. Here, we have compared these two approaches by means of computer simulation, using 448 parameter sets, including evolutionary rate, sequence length, base composition, and transition/transversion rate bias. In these simulations, we emphasized a worst-case scenario analysis in which 100 replicate datasets for each evolutionary parameter set (gene) were generated, and the replicate dataset that produced a tree topology showing the largest number of phylogenetic errors was selected to represent that parameter set. Both randomly selected and worst-case replicates were utilized to compare the consensus and concatenation approaches primarily using the neighbor-joining (NJ) method. We find that the concatenation approach yields more accurate trees, even when the sequences concatenated have evolved with very different substitution patterns and no attempts are made to accommodate these differences while inferring phylogenies. These results appear to hold true for parsimony and likelihood methods as well. The concatenation approach shows >95% accuracy with only 10 genes. However, this gain in accuracy is sometimes accompanied by reinforcement of certain systematic biases, resulting in spuriously high bootstrap support for incorrect partitions, whether we employ site, gene, or a combined bootstrap resampling approach. Therefore, it will be prudent to report the number of individual genes supporting an inferred clade in the concatenated sequence tree, in addition to the bootstrap support.

Animals↗

Do large dogs die young?

In most animal taxa, longevity increases with body size across species, as predicted by the oxidative stress theory of aging. In contrast, in within-species comparisons of mammals and especially domestic dogs (e.g. Patronek et al., '97; Michell, '99; Egenvall et al., 2000; Speakman et al., 2003), longevity decreases with body size. We explore two datasets for dogs and find support for a negative relationship between size and longevity if we consider variation across breeds. Within breeds, however, the relationship is not negative and is slightly, but significantly, positive in the larger of the two datasets. The negative across-breed relationship is probably the consequence of short life spans in large breeds. Artificial selection for extremely high growth rates in large breeds appears to have led to developmental diseases that seriously diminish longevity.

Animals↗

MELD-XI: a rational approach to "sickest first" liver transplantation in cirrhotic patients requiring anticoagulant therapy.

Priority for "sickest first" liver transplantation (LT) in the United States is determined by the model for end-stage liver disease (MELD). MELD is a good predictor of short-term mortality in cirrhosis, but it can overestimate risk when international normalized ratio (INR) is artificially elevated by anticoagulation. An alternate prognostic index omitting INR is needed in this situation. We retrospectively analyzed survival data for 554 cirrhotic veterans referred for consideration of LT prior to December 1, 2003 (training group). Using logistic regression we derived a predictive formula for 90-day pretransplant mortality incorporating bilirubin and creatinine but omitting INR. We normalized this formula to the same scale as MELD using linear regression. This yielded MELD-XI (for MELD excluding INR) = 5.11 Ln(bilirubin) + 11.76 Ln(creatinine) + 9.44. Accuracy of MELD-XI was validated in a holdout group of 278 cirrhotic veterans referred after December 1, 2003, and in an independent validation dataset of 7,203 cirrhotic adults listed for LT in the United States between May 1, 2001, and October 31, 2001. MELD-XI and MELD correlated well in training, holdout, and independent validation cohorts (r = 0.930, 0.954, and 0.902, respectively). In the holdout cohort, c-statistics of MELD vs. MELD-XI for mortality were, respectively, 0.939 vs. 0.906 at 30 days;0.860 vs. 0.841 at 60 days; 0.842 vs. 0.829 at 90 days; and 0.795 vs. 0.797 at 180 days. In the independent validation dataset, c-statistics for MELD vs. MELD-XI as predictors of 90-day survival were, respectively, 0.857 vs. 0.843 in noncholestatic liver diseases and 0.905 vs. 0.894 in cholestatic liver diseases. Comparable MELD and MELD-XI scores were associated with comparable prognosis. In conclusion, MELD-XI, despite omission of INR, is nearly as accurate as MELD in predicting short-term survival in cirrhosis. In patients treated with oral anticoagulants, substitution of MELD-XI for MELD may permit more accurate assessment of risk and more rational assignment of "sickest first" priority for LT.

Adult↗

Bioinformatics Analysis and Experimental Validation of Key Genes Associated With Hypoxia and Ischemia in Myocardial Infarction.

BACKGROUND: This study aimed to screen and identify core hypoxia-ischemia-related genes associated with myocardial infarction (MI). METHOD: Two transcriptomic datasets, GSE97320 and GSE48060, were retrieved from the Gene Expression Omnibus (GEO) database. After data integration and batch effect elimination, differential expression analysis was performed to screen differentially expressed genes (DEGs), and the corresponding visualization analysis was conducted. Hypoxia-ischemia-related genes were acquired from the GeneCards database; hypoxia-ischemia related genes (HIRGs) were subsequently identified by intersecting the retrieved genes with screened DEGs. Gene Ontology (GO) functional enrichment and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analyses were implemented to explore the biological functions and underlying signaling pathways of HIRGs. A combination of protein-protein interaction (PPI) network analysis and random forest (RF) algorithm was applied to screen hub genes from HIRGs. The external GEO dataset GSE66360 was utilized to validate the expression patterns of candidate hub genes. Furthermore, an acute myocardial infarction (AMI) mouse model was established, and quantitative real-time polymerase chain reaction (qPCR) was performed to detect the mRNA expression levels of hub genes in myocardial tissues for in&#xa0;vivo validation. RESULTS: A total of 633 DEGs and 308 hypoxia-ischemia-related genes were screened in the present study, among which 21 overlapping HIRGs were obtained. PLAUR and IL1B were finally identified as two hub genes from HIRGs based on PPI network and random forest algorithm. The qPCR results revealed that the expression levels of PLAUR and IL1B were significantly upregulated in the AMI group compared with the sham operation group (p&#x2009;<&#x2009;0.05). CONCLUSION: The present findings demonstrated that PLAUR and IL1B serve as pivotal genes involved in the pathological hypoxia-ischemia process of AMI. These two genes may act as novel biomarkers and promising therapeutic targets for the recognition and clinical intervention of hypoxia-ischemia injury following AMI.

Myocardial Infarction↗

Estimation of relative cardiovascular pressures using time-resolved three-dimensional phase contrast MRI.

Accurate, easy-to-use, noninvasive cardiovascular pressure registration would be an important addition to the diagnostic armamentarium for assessment of cardiac function. A novel noninvasive and three-dimensional (3D) technique for estimation of relative cardiovascular pressures is presented. The relative pressure is calculated using the Navier-Stokes equations along user-defined lines placed within a time-resolved 3D phase contrast MRI dataset. The lines may be either straight or curved to follow an actual streamline. The technique is validated in an in vitro model and tested on in vivo cases of normal and abnormal transmitral pressure differences and intraaortic flow. The method supplements an intuitive visualization technique for cardiovascular flow, 3D particle trace visualization, with a quantifiable diagnostic parameter estimated from the same dataset.

Adult↗

Simultaneous acquisition of multiple resolution images for dynamic contrast enhanced imaging of the breast.

An imaging technique is described that allows the reconstruction of a series of images at high temporal rates, while simultaneously providing images at high spatial resolution. The method allows one to arbitrarily choose from among several combinations of temporal/spatial resolutions during postprocessing. This flexibility is accomplished by strategically interleaving multiple undersampled projection reconstruction datasets (or subapertures), in which each set can be used to reconstruct a high temporal resolution image. Images with increasingly higher spatial resolutions can subsequently be formed by combining two or more subaperture datasets. The technique is demonstrated in vivo to assess the kinetics of contrast enhancement and to visualize the architectural features of suspicious breast lesions.

Aged↗

Accelerating cardiac cine 3D imaging using k-t BLAST.

By exploiting spatiotemporal correlations in cardiac acquisitions using k-t BLAST, gated cine 3D acquisitions of the heart were accelerated by a net factor of 4.3, making single breathhold acquisitions possible. Sparse sampling of k-t space along a sheared grid pattern was implemented into a cine 3D SSFP sequence. The acquisition of low-resolution training data, which was required to resolve aliasing in the k-t BLAST method, was either interleaved into the sampling process or obtained in a separate prescan to allow for shorter breathhold durations in patients with heart disease. Volumetric datasets covering the heart with 20 slices at a spatial resolution of 2 x 2 x 5 mm3 were recorded with 20 cardiac phases in a total breathhold duration of 25-27 sec, or 18 sec if partial Fourier sampling was additionally employed. The feasibility of the method was demonstrated on healthy volunteers and on patients. The comparison of endocardial area derived from single slices of the 3D dataset with values extracted from separate single-slice acquisitions showed no significant differences. By shortening the acquisition substantially, k-t BLAST may greatly facilitate volumetric imaging of the heart for evaluation of regional wall motion and the assessment of ventricular volume and ejection fraction.

Computer Simulation↗

Detection and volume determination of colonic tumors in Min mice by magnetic resonance micro-imaging.

We applied MRI to the in vivo detection of spontaneous colorectal tumors in a unique mouse model, the Fox Chase Cancer Center (FCCC) ApcMIN mouse. Unlike other Min (multiple intestinal neoplasia) strains, FCCC ApcMIN animals develop an appreciable number of tumors in the large intestine, which makes them an appropriate mouse model for colon cancer in humans. We describe a method for marking the colon on MRI data sets that involves a bowel-cleansing procedure and the insertion of a polyurethane tube (filled with an MRI contrast agent) fully into the colon. We found that tumors as small as 1.5 mm in diameter can be consistently identified from MRI datasets with a voxel size of 0.1 mm x 0.133 mm x 0.133 mm. Tumor volumes were determined from the MRM data sets with the use of a novel approach to planimetry in 3D data sets. We observed a correlation between tumor volume (as measured from the MRI datasets) and tumor weight of 0.942, and a P-value of 0.008, based on Spearman's test. These data show that MRI can be used to accurately monitor tumor growth in mouse models of colorectal carcinogenesis.

Animals↗

Accelerating MRI by skipped phase encoding and edge deghosting (SPEED).

A fast imaging method called skipped phase encoding and edge deghosting (SPEED) is introduced. The k-space is sparsely sampled into three interleaved datasets, each with a skip-size N and a relative shift in phase encoding (PE). These datasets are separately reconstructed by 2DFT and edge-enhanced by a differential filter in the PE direction, resulting in edge maps with phase-shifted aliasing ghosts. The sparseness of edges reduces the chance of ghost overlapping. Typical ghosted-edge maps can be adequately modeled with only two dominating ghost layers that are resolved from a set of three equations using least-square error minimization, yielding N ghost maps of different orders that can be registered and averaged into a single deghosted-edge map for noise and artifact reduction. Finally, the deghosted-edge map is transformed into a deghosted image by an inverse filter. A few central k-space lines are collected without PE skip to aid the inverse filtering. SPEED has been demonstrated by in vivo data to reduce scan time considerably without noticeable artifacts. It has various potential applications, such as MR angiography (MRA), where the signal itself is sparse. As an independent method, SPEED can be combined with other fast imaging methods for further acceleration.

Algorithms↗

Application of partial least squares discriminant analysis to two-dimensional difference gel studies in expression proteomics.

Two-dimensional difference gel electrophoresis (DIGE) is a tool for measuring changes in protein expression between samples involving pre-electrophoretic labeling ith cyanine dyes. In multi-gel experiments, univariate statistical tests have been used to identify differential expression between sample types by looking for significant changes in spot volume. Multivariate statistical tests, which look for correlated changes between sample types, provide an alternate approach for identifying spots with differential expression. Partial least squares-discriminant analysis (PLS-DA), a multivariate statistical approach, was combined with an iterative threshold process to identify which protein spots had the greatest contribution to the model, and compared to univariate test for three datasets. This included one dataset where no biological difference was expected. The novel multivariate approach, detailed here, represents a method to complement the univariate approach in identification of differentially expressed protein spots. This new approach has the advantages of reduced risk of false-positives and the identification of spots that are significantly altered in terms of correlated expression rather than absolute expression values.

Analysis of Variance↗

Protein and peptide identification algorithms using MS for use in high-throughput, automated pipelines.

Current proteomics experiments can generate vast quantities of data very quickly, but this has not been matched by data analysis capabilities. Although there have been a number of recent reviews covering various aspects of peptide and protein identification methods using MS, comparisons of which methods are either the most appropriate for, or the most effective at, their proposed tasks are not readily available. As the need for high-throughput, automated peptide and protein identification systems increases, the creators of such pipelines need to be able to choose algorithms that are going to perform well both in terms of accuracy and computational efficiency. This article therefore provides a review of the currently available core algorithms for PMF, database searching using MS/MS, sequence tag searches and de novo sequencing. We also assess the relative performances of a number of these algorithms. As there is limited reporting of such information in the literature, we conclude that there is a need for the adoption of a system of standardised reporting on the performance of new peptide and protein identification algorithms, based upon freely available datasets. We go on to present our initial suggestions for the format and content of these datasets.

Algorithms↗

Protein probabilities in shotgun proteomics: evaluating different estimation methods using a semi-random sampling model.

The calculation of protein probabilities is one of the most intractable problems in large-scale proteomic research. Current available estimating methods, for example, ProteinProphet, PROT_PROBE, Poisson model and two-peptide hits, employ different models trying to resolve this problem. Until now, no efficient method is used for comparative evaluation of the above methods in large-scale datasets. In order to evaluate these various methods, we developed a semi-random sampling model to simulate large-scale proteomic data. In this model, the identified peptides were sampled from the designed proteins and their cross-correlation scores were simulated according to the results from reverse database searching. The simulated result of 18 control proteins was consistent with the experimental one, demonstrating the efficiency of our model. According to the simulated results of human liver sample, ProteinProphet returned slightly higher probabilities and lower specificity than real cases. PROT_PROBE was a more efficient method with higher specificity. Predicted results from a Poisson model roughly coincide with real datasets, and the method of two-peptide hits seems solid but imprecise. However, the probabilities of identified proteins are strongly correlated with several experimental factors including spectra number, database size and protein abundance distribution.

Chromatography, Liquid↗

Searching for biomarkers of heart failure in the mass spectra of blood plasma.

We have developed a technique for analysing blood plasma using MALDI-MS with subsequent data analysis to identify significant and specific differences between heart failure (HF) patients and healthy individuals. A training dataset comprising 100 HF patients and 100 healthy individuals was used to search for biomarkers (m/z range 1000-10,000). EWP cartridges when used in tandem with microcon centrifugal filters were found to give the best results. A data management chain including event binning, background subtraction and feature extraction was developed to reduce the data, and statistical analysis was used to map feature intensities on to a common scale. Various mathematical approaches including a simple cumulative score, support vector machines (SVM) and genetic algorithms (GAs) were then used to combine the results from individual features and provide a robust classification algorithm. The SVM gave the most promising results (accuracy 95%, receiver operating characteristic (ROC) score of 0.997 using 18 selected features). Finally, a test dataset comprising a further 32 HF patients and 20 controls was used to verify that the 18 putative biomarkers and classification algorithms gave reliable predictions (accuracy 88.5%, ROC score 0.998).

Biomarkers↗

Structural motifs at protein-protein interfaces: protein cores versus two-state and three-state model complexes.

The general similarity in the forces governing protein folding and protein-protein associations has led us to examine the similarity in the architectural motifs between the interfaces and the monomers. We have carried out extensive, all-against-all structural comparisons between the single-chain protein structural dataset and the interface dataset, derived both from all protein-protein complexes in the structural database and from interfaces generated via an automated crystal symmetry operation. We show that despite the absence of chain connections, the global features of the architectural motifs, present in monomers, recur in the interfaces, a reflection of the limited set of the folding patterns. However, although similarity has been observed, the details of the architectural motifs vary. In particular, the extent of the similarity correlates with the consideration of how the interface has been formed. Interfaces derived from two-state model complexes, where the chains fold cooperatively, display a considerable similarity to architectures in protein cores, as judged by the quality of their geometric superposition. On the other hand, the three-state model interfaces, representing binding of already folded molecules, manifest a larger variability and resemble the monomer architecture only in general outline. The origin of the difference between the monomers and the three-state model interfaces can be understood in terms of the different nature of the folding and the binding that are involved. Whereas in the former all degrees of freedom are available to the backbone to maximize favorable interactions, in rigid body, three-state model binding, only six degrees of freedom are allowed. Hence, residue or atom pair-wise potentials derived from protein-protein associations are expected to be less accurate, substantially increasing the number of computationally acceptable alternate binding modes (Finkelstein et al., 1995).

Chemical Phenomena↗

Analysis of zinc binding sites in protein crystal structures.

The geometrical properties of zinc binding sites in a dataset of high quality protein crystal structures deposited in the Protein Data Bank have been examined to identify important differences between zinc sites that are directly involved in catalysis and those that play a structural role. Coordination angles in the zinc primary coordination sphere are compared with ideal values for each coordination geometry, and zinc coordination distances are compared with those in small zinc complexes from the Cambridge Structural Database as a guide of expected trends. We find that distances and angles in the primary coordination sphere are in general close to the expected (or ideal) values. Deviations occur primarily for oxygen coordinating atoms and are found to be mainly due to H-bonding of the oxygen coordinating ligand to protein residues, bidentate binding arrangements, and multi-zinc sites. We find that H-bonding of oxygen containing residues (or water) to zinc bound histidines is almost universal in our dataset and defines the elec-His-Zn motif. Analysis of the stereochemistry shows that carboxyl elec-His-Zn motifs are geometrically rigid, while water elec-His-Zn motifs show the most geometrical variation. As catalytic motifs have a higher proportion of carboxyl elec atoms than structural motifs, they provide a more rigid framework for zinc binding. This is understood biologically, as a small distortion in the zinc position in an enzyme can have serious consequences on the enzymatic reaction. We also analyze the sequence pattern of the zinc ligands and residues that provide elecs, and identify conserved hydrophobic residues in the endopeptidases that also appear to contribute to stabilizing the catalytic zinc site. A zinc binding template in protein crystal structures is derived from these observations.

Aspartic Acid↗

Genomic scan of 254 hereditary prostate cancer families.

Hereditary prostate cancer (HPC) is a genetically heterogeneous disease, complicating efforts to map and clone susceptibility loci. We have used stratification of a large dataset of 254 HPC families in an effort to improve power to detect HPC loci and to understand what types of family features may improve locus identification. The strongest result is that of a dominant locus at 6p22.3 (heterogeneity LOD (HLOD) = 2.51), the evidence for which is increased by consideration of the age of PC onset (HLOD = 3.43 in 214 families with median age-of-onset 56-72 years) and co-occurrence of primary brain cancer (HLOD = 2.34 in 21 families) in the families. Additional regions for which we observe modest evidence for linkage include chromosome 7q and 17p. Only weak evidence of several previously implicated HPC regions is detected. These analyses support the existence of multiple HPC loci, whose presence may be best identified by analyses of large, including pooled, datasets which consider locus heterogeneity.

Aged↗