Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Pollution regimes and variability in river water quality across the Humber catchment: interrogation and mapping of an extensive and highly heterogeneous spatial dataset

Water quality inter-relationships based on an extensive Environment Agency database are used to identify seven key types of clean and polluted riverine environments within the Humber catchment. These types were selected on the basis of linear separations between average concentrations of ammonium, chloride, dissolved nickel and chromium and acid-available particulate nickel and chromium. However, a more detailed analysis of determinand and flow relationships at each type locality revealed complex patterns due to the highly variable nature of pollutant sources at the local scale. The information presented contrasts systems at a larger scale, where simpler integrated features are observed, and which can be represented by two-component mixing models. Scale and hydrology seem to be important features in determining the simplicity or complexity of determinand inter-relationships, and the implications of these findings are discussed.

Journal Article↗

Analytical methods for detecting paralogy in molecular datasets.

Paralogy (common ancestry through gene duplication rather than speciation) is widely recognized as an important problem for molecular systematists. This chapter introduces the concepts of paralogy and orthology and explains why paralogy can complicate both systematic work and other studies of molecular evolution. The definition of paralogy is explicitly phylogenetic, and phylogenetic methods are crucial in elucidating the pattern of paralogy. In particular, knowledge of the species phylogeny is key. I introduce the theory behind methods for detecting paralogy and briefly discuss two particular software implementations of phylogenetic methods to detect paralogy from molecular data. I also introduce a statistical method for detecting paralogy and some future directions for work on paralogy detection.

Animals↗

Origin of Macaronesian Sideritis L. (Lamioideae: Lamiaceae) inferred from nuclear and chloroplast sequence datasets.

Sideritis L. (Lamiaceae) comprises approximately 150 species of annuals and perennials distributed chiefly in the Mediterranean region. The majority of the species belong to the continental subgenus Sideritis which is divided into two perennial (Sideritis and Empedoclea) and two annual (Hesiodia and Burgsdorfia) sections. Twenty-three species are woody perennials endemic to the Macaronesian archipelagos of Madeira and the Canary Islands. In an effort to determine the continental origin of the insular group, we constructed independent phylogenies comprising sequence data from both chloroplast and nuclear markers. Sampling included 7 island taxa drawn from the Macaronesian subgenus Marrubiastrum and 25 continental taxa representing all four sections of subgenus Sideritis. Subgenus Marrubiastrum and the two continental perennial sections form well-supported monophyletic groups in both individual and combined analyses. The annual sections are not monophyletic in any analysis; further sampling of annual taxa is needed to resolve these relationships. All analyses identified Sideritis cossoniana, an annual species from Morocco, as the closest continental relative of the Macaronesian group. This contrasts with the hypothesis of earlier workers who suggested that the insular taxa were most closely related to eastern Mediterranean species of the genus. The phylogenies also demonstrate a distinct increase in woodiness among the Macaronesian species relative to their continental congeners, providing further support for the secondary nature of woodiness in island plants.

Biological Evolution↗

Publishing large proteome datasets: scientific policy meets emerging technologies.

Currently, there are various approaches to proteomic analyses based on either 2D gel or HPLC separation platforms, generating data of different formats, structures and types. Identification of these separated proteins or peptide fragments is typically achieved by mass spectrometry (MS) measurements that use either accurate mass measurements or fragmentation (MS-MS) information. Integrating the information generated from these different platforms is essential if proteomics is to succeed. A further challenge lies in generating standards that can accept the hundreds-of-thousands of mass spectra produced per analysis based on threshold or probability measurements. Finally, peer review and electronic publication processes will be crucial to the dissemination and use of proteomic information. Merging the policy requirements of data-intensive research with information technology will enable scientists to gain real value from global proteomics information.

Chromatography, Liquid↗

Application of fast Fourier transform cross-correlation for the alignment of large chromatographic and spectral datasets.

Preprocessing of chromatographic and spectral data is an important aspect of analytical sciences. In particular, recent advances in proteomics have resulted in the generation of large data sets that require analysis. To assist accurate comparison of chemical signals, we propose two methods for the alignment of multiple spectral data sets. Based on methods previously described, each chromatograph or spectrum to be aligned is divided and aligned as individual segments to a reference. However, our methods make use of fast Fourier transform for the rapid computation of a cross-correlation function that enables alignments between samples to be optimized. The proposed methods are demonstrated in comparison with an existing method on a chromatographic and a mass spectral data set. It is shown that our methods provide an advantage of speed and a reduction of the number of input parameters required. The software implementations for the proposed alignment methods are available under the downloads section at http://ptcl.chem.ox.ac.uk/~jwong/specalign.

Algorithms↗

G-protein-coupled receptor affinity prediction based on the use of a profiling dataset: QSAR design, synthesis, and experimental validation.

A QSAR model accounting for "average" G-protein-coupled receptor (GPCR) binding was built from a large set of experimental standardized binding data (1939 compounds systematically tested over 40 different GPCRs) and applied to the design of a library of "GPCR-predicted" compounds. Three hundred and sixty of these compounds were randomly selected and tested in 21 GPCR binding assays. Positives were defined by their ability to inhibit by more than 70% the binding of reference compounds at 10 microM. A 5.5-fold enrichment in positives was observed when comparing the "GPCR-predicted" compounds with 600 randomly selected compounds predicted as "non-GPCR" from a general collection. The model was efficient in predicting strongest binders, since enrichment was greater for higher cutoffs. Significant enrichment was also observed for peptidic GPCRs and receptors not included to develop the QSAR model, suggesting the usefulness of the model to design ligands binding with newly identified GPCRs, including orphan ones.

Combinatorial Chemistry Techniques↗

Integration of genomic datasets to predict protein complexes in yeast.

The ultimate goal of functional genomics is to define the function of all the genes in the genome of an organism. A large body of information of the biological roles of genes has been accumulated and aggregated in the past decades of research, both from traditional experiments detailing the role of individual genes and proteins, and from newer experimental strategies that aim to characterize gene function on a genomic scale. It is clear that the goal of functional genomics can only be achieved by integrating information and data sources from the variety of these different experiments. Integration of different data is thus an important challenge for bioinformatics. The integration of different data sources often helps to uncover non-obvious relationships between genes, but there are also two further benefits. First, it is likely that whenever information from multiple independent sources agrees, it should be more valid and reliable. Secondly, by looking at the union of multiple sources, one can cover larger parts of the genome. This is obvious for integrating results from multiple single gene or protein experiments, but also necessary for many of the results from genome-wide experiments since they are often confined to certain (although sizable) subsets of the genome. In this paper, we explore an example of such a data integration procedure. We focus on the prediction of membership in protein complexes for individual genes. For this, we recruit six different data sources that include expression profiles, interaction data, essentiality and localization information. Each of these data sources individually contains some weakly predictive information with respect to protein complexes, but we show how this prediction can be improved by combining all of them. Supplementary information is available at http:// bioinfo.mbb.yale.edu/integrate/interactions/.

Cell Cycle↗

Dubious dataset.

Explore the source record for details and available documents.

Bias↗

Locational uncertainty in georeferencing public health datasets.

The assignment of locational attributes to a study subject in epidemiologic analyses is commonly referred to as georeferencing. When georeferencing study subjects to a point location using their residential street address, most researchers rely on the street centerline data model. This study assessed the potential locational bias introduced using street centerline data. It also evaluated georeferencing effects on a location-dependent, exposure assessment process. For comparison purposes, subjects were georeferenced to the center of their residential parcel of land using digitized parcel maps. A total of 10,026 study subjects residing in Jefferson County, Alabama were georeferenced using both street centerline and residential parcel methods. The mean nondirectional, linear distance between points georeferenced using both methods was 246 ft with a range of 11 to 13,260 ft. Correlation coefficients comparing differences in exposure estimates were generated for all 10,026 subjects. Coefficients increased as the geographic areas of analysis around study subjects increased, indicating the influence of nondifferential exposure misclassification.

Bias↗

Automated quantification and reconstruction of collagen matrix from 3D confocal datasets.

The geometrical structure of fibrous extracellular matrix (ECM) impacts on its biological function. In this report, we demonstrate a new algorithm designed to extract quantitative structural information about individual collagen fibres (orientation, length and diameter) from 3D backscattered-light confocal images of collagen gels. The computed quantitative data allowed us to create surface-rendered 3D images of the investigated sample.

Animals↗

Statistical analysis of the effect of high dilutions of arsenic in a large dataset from a wheat germination model.

This paper describes the statistical analysis of a series of experiments using a simple biological model (wheat germination in vitro), where a large number of wheat seeds were treated with homeopathic potencies of Arsenic trioxide. Some potencies, such as As2O3 40x, 42x and 45x, have repeatedly shown a significant stimulating effect on germination compared to controls, whereas As2O3 35x has a significant inhibiting effect. In some experiments the seeds were stressed before the experiment with a sublethal dose of the same substance. We performed a statistical analysis, both for stressed and non-stressed seed groups, using Poisson distribution as a suitable model for representing the number of non-germinated seeds in a standard experiment with 33 seeds in the same Petri dish. Finally, we have considered the most repeated potencies (30x and 45x), computing the sample odds ratio (OR) and a 95% confidence interval (CI) for the population OR. Our results show significant reproducible effects of some As2O3 decimal potencies, particularly As2O3 45x. In stressed seeds, even decimal potencies of water seem to give significant results compared to control, whereas high dilutions of As2O3 without potentization never show significant effects.

Arsenic Trioxide↗

A dataset of human liver proteins identified by protein profiling via isotope-coded affinity tag (ICAT) and tandem mass spectrometry.

Proteins from human liver carcinoma Huh7 cells, representing transformed liver cells, and cultured primary human fetal hepatocytes (HFH) and human HH4 hepatocytes, representing nontransformed liver cells, were extracted and processed for proteome analysis. Proteins from stimulated cells (interferon-alpha treatment for the Huh7 and HFH cells and induction of hepatitis C virus [HCV] proteins for the HH4 cells) and corresponding control cells were labeled with light and heavy cleavable ICAT reagents, respectively. The labeled samples were combined, trypsinized, and subject to cation-exchange and avidin-affinity chromatographies. The resulting cysteine-containing peptides were analyzed by microcapillary LC-MS/MS. The MS/MS spectra were initially analyzed by searching the human International Protein Index database using the SEQUEST software (1). Subsequently, new statistical algorithms were applied to the collective SEQUEST search results of each experiment. First, the PeptideProphet software (2) was applied to discriminate true assignments of MS/MS spectra to peptide sequences from false assignments, to assign a probability value for each identified peptide, and to compute the sensitivity and error rate for the assignment of spectra to sequences in each experiment. Second, the ProteinProphet software (3) was used to infer the protein identifications and to compute probabilities that a protein had been correctly identified, based on the available peptide sequence evidence. The resulting protein lists were filtered by a ProteinProphet probability score p > or = 0.5, which corresponded to an error rate of less than 5%. A total of 1,296, 1,430, and 1,476 proteins or related protein groups were identified in three subdatasets from the Huh7, HFH, and HH4 cells, respectively. In total, these subdatasets contained 2,486 unique protein identifications from human liver cells. An increase of the threshold to p > or = 0.9 (corresponding to an error rate of less than 1%) resulted in 2,159 unique protein identifications (1,146, 1,235, and 1,318 for the Huh7, HFH, and HH4 cells, respectively).

Algorithms↗

A dataset of human cornea proteins identified by Peptide mass fingerprinting and tandem mass spectrometry.

Diseases of the cornea are extremely common and cause severe visual impairment worldwide. To explore the basic molecular mechanisms involved in corneal health and disease, the present study characterizes the proteome of the normal human cornea. All proteins were extracted from the central 7-mm region of 12 normal human donor corneas containing all layers: epithelium, Bowman's layer, stroma, Descemet's membrane, and endothelium. Proteins were fractionated and identified using two different procedures: (i) two-dimensional gel electrophoresis and protein identification by MALDI-MS and (ii) strong cation exchange or one-dimensional SDS gel electrophoresis followed by LC-MS/MS. All together, 141 distinct proteins were identified of which 99 had not previously been identified in any mammalian corneas by direct protein identification methods. The characterized proteins are involved in many processes including antiangiogenesis, antimicrobial defense, protection from and transport of heme and iron, tissue protection against UV radiation and oxidative stress, cell metabolism, and maintenance of intracellular and extracellular structures and stability. This proteome study of the healthy human cornea provides a basis for further analysis of corneal diseases and the design of bioengineered corneas.

Chemical Fractionation↗

EBP, a program for protein identification using multiple tandem mass spectrometry datasets.

MS/MS combined with database search methods can identify the proteins present in complex mixtures. High throughput methods that infer probable peptide sequences from enzymatically digested protein samples create a challenge in how best to aggregate the evidence for candidate proteins. Typically the results of multiple technical and/or biological replicate experiments must be combined to maximize sensitivity. We present a statistical method for estimating probabilities of protein expression that integrates peptide sequence identifications from multiple search algorithms and replicate experimental runs. The method was applied to create a repository of 797 non-homologous zebrafish (Danio rerio) proteins, at an empirically validated false identification rate under 1%, as a resource for the development of targeted quantitative proteomics assays. We have implemented this statistical method as an analytic module that can be integrated with an existing suite of open-source proteomics software.

Algorithms↗