Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Datasets as Topic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

yMGV: a database for visualization and data mining of published genome-wide yeast expression data.

The yeast Microarray Global Viewer (yMGV) is an on-line database providing a synthetic view of the transcriptional expression profiles of Saccharomyces cerevisiae genes in most of the published expression datasets. yMGV displays a one-screen graphical representation of gene expression variations for each published genome-wide experiment, allowing quick retrieval of experimental conditions affecting expression of this gene. yMGV also provides tools to isolate groups of genes sharing similar transcription profiles in a defined subset of experiments. Additionally, yMGV furnishes a set of statistical tools for critical assessment of published data. We therefore believe that yMGV is an efficient tool that affords a quick and comprehensive overview of microarray data and generates new gene classifications. As of 20 March 2001 the yMGV database contains 6 000 000 measurements, representing genome-wide expression comparisons of 932 experiments from 39 microarray publications. The yMGV interface is available at http://transcriptome.ens.fr/ymgv/.

Computational Biology↗

Identification of susceptibility loci for complex diseases in a case-control association study using the Genetic Analysis Workshop 14 dataset.

Although current methods in genetic epidemiology have been extremely successful in identifying genetic loci responsible for Mendelian traits, most common diseases do not follow simple Mendelian modes of inheritance. It is important to consider how our current methodologies function in the realm of complex diseases. The aim of this study was to determine the ability of conventional association methods to fine map a locus of interest. Six study populations were selected from 10 replicates (New York) from the Genetic Analysis Workshop 14 simulated dataset and analyzed for association between the disease trait and locus D2. Genotypes from 45 single-nucleotide polymorphisms in the telomeric region of chromosome 3 were analyzed by Pearson's chi-square tests for independence to test for association with the disease trait of interest. A significant association was detected within the region; however, it was found 3 cM from the documented location of the D2 disease locus. This result was most likely due to the method used for data simulation. In general, this study showed that conventional case-control association methods could detect disease loci responsible for the development of complex traits.

Case-Control Studies↗

Clinical quality needs complex adaptive systems and machine learning.

The vast increase in clinical data has the potential to bring about large improvements in clinical quality and other aspects of healthcare delivery. However, such benefits do not come without cost. The analysis of such large datasets, particularly where the data may have to be merged from several sources and may be noisy and incomplete, is a challenging task. Furthermore, the introduction of clinical changes is a cyclical task, meaning that the processes under examination operate in an environment that is not static. We suggest that traditional methods of analysis are unsuitable for the task, and identify complexity theory and machine learning as areas that have the potential to facilitate the examination of clinical quality. By its nature the field of complex adaptive systems deals with environments that change because of the interactions that have occurred in the past. We draw parallels between health informatics and bioinformatics, which has already started to successfully use machine learning methods.

Artificial Intelligence↗

Identification of tryptic peptides from large databases using multiplexed tandem mass spectrometry: simulations and experimental results.

Multiplexed tandem mass spectrometry (MS/MS) has recently been demonstrated as a means to increase the throughput of peptide identification in liquid chromatography (LC) MS/MS experiments. In this approach, a set of parent species is dissociated simultaneously and measured in a single spectrum (in the same manner that a single parent ion is conventionally studied), providing a gain in sensitivity and throughput proportional to the number of species that can be simultaneously addressed. In the present work, simulations performed using the Caenorhabditis elegans predicted proteins database show that multiplexed MS/MS data allow the identification of tryptic peptides from mixtures of up to ten peptides from a single dataset with only three "y" or "b" fragments per peptide and a mass accuracy of 2.5 to 5 ppm. At this level of database and data complexity, 98% of the 500 peptides considered in the simulation were correctly identified. This compares favorably with the rates obtained for classical MS/MS at more modest mass measurement accuracy. LC multiplexed Fourier transform-ion cyclotron resonance MS/MS data obtained from a 66 kDa protein (bovine serum albumin) tryptic digest sample are presented to illustrate the approach, and confirm that peptides can be effectively identified from the C. elegans database to which the protein sequence had been appended.

Algorithms↗

Identifying residues in natural organic matter through spectral prediction and pattern matching of 2D NMR datasets.

This paper describes procedures for the generation of 2D NMR databases containing spectra predicted from chemical structures. These databases allow flexible searching via chemical structure, substructure or similarity of structure as well as spectral features. In this paper we use the biopolymer lignin as an example. Lignin is an important and relatively recalcitrant structural biopolymer present in the majority of plant biomass. We demonstrate how an accurate 2D NMR database of approximately 600 2D spectra of lignin fragments can be easily constructed, in approximately 2 days, and then subsequently show how some of these fragments can be identified in soil extracts through the use of various search tools and pattern recognition techniques. We demonstrate that once identified in one sample, similar residues are easily determined in other soil extracts. In theory, such an approach can be used for the analysis of any organic mixtures.

Benzopyrans↗

Computational assignment of the EC numbers for genomic-scale analysis of enzymatic reactions.

The EC (Enzyme Commission) numbers represent a hierarchical classification of enzymatic reactions, but they are also commonly utilized as identifiers of enzymes or enzyme genes in the analysis of complete genomes. This duality of the EC numbers makes it possible to link the genomic repertoire of enzyme genes to the chemical repertoire of metabolic pathways, the process called metabolic reconstruction. Unfortunately, there are numerous reactions known to be present in various pathways, but they will never get EC numbers because the EC number assignment requires published articles on full characterization of enzymes. Here we report a computerized method to automatically assign the EC numbers up to the sub-subclasses, i.e., without the fourth serial number for substrate specificity, given pairs of substrates and products. The method is based on a new classification scheme of enzymatic reactions, named the RC (reaction classification) number. Each reaction in the current dataset of the EC numbers is first decomposed into reactant pairs. Each pair is then structurally aligned to identify the reaction center, the matched region, and the difference region. The RC number represents the conversion patterns of atom types in these three regions. We examined the correspondence between computationally assigned RC numbers and manually assigned EC numbers by the jackknife cross-validation test and found that the EC sub-subclasses could be assigned with the accuracy of about 90%. Furthermore, we examined the correlation with genomic information as represented by the KEGG ortholog clusters (OC) and confirmed that the RC numbers are correlated not only with elementary reaction mechanisms but also with protein families.

Computers↗

Analysis of the CD3 gene region and type 1 diabetes: application of fluorescence-based technology to linkage disequilibrium mapping.

The CD3 gene region on chromosome 11q23 has been implicated in susceptibility to type 1 (insulin-dependent) diabetes mellitus. Using semi-automated fluorescence-based technology, we have undertaken association and linkage analysis of a dinucleotide microsatellite in the CD3 delta (CD3D) gene. We have also performed a large case-control analysis of a restriction fragment length polymorphism (RFLP) in the CD3 epsilon (CD3E) gene, 26 kb from CD3D. We found no evidence for the previously reported association between the 8 kb allele of the RFLP and disease in a UK dataset of 403 diabetic patients and 446 nondiabetic controls. Furthermore, the use of the transmission/disequilibrium test (TDT) showed no evidence of linkage or association to type 1 diabetes at either marker locus. We conclude that the CD3 gene region does not contribute significantly to IDDM susceptibility. We have successfully applied semi-automated, fluorescence-based technology to undertake association analysis on the CD3D microsatellite. Moreover, by analysing 94 other dinucleotide repeat markers, we conclude that fluorescence-based methodology can generally be applied to large-scale, semi-automated association studies with most microsatellite markers.

Adolescent↗

Detection of differentially expressed proteins in early-stage melanoma patients using SELDI-TOF mass spectrometry.

Tumor progression is a dynamic sequence of events that involves specific protein changes. We hypothesized that Surface Enhanced Laser Desorption/Ionization (SELDI) mass spectrometric analysis of sera from patients with AJCC stage I and II melanoma with negative loco-regional lymph nodes could identify potential melanoma-associated protein biomarkers of disease recurrence. Serum specimens were collected from 49 patients who developed recurrence (n = 25) or remained free of recurrence (n = 24) without evidence of disease following complete resection (AJCC stage I and II). Follow-up was longer than 5 years. Serum proteins were denatured and applied onto two protein chip chemistry surfaces (weak cationic WCX2; metal-binding, IMAC3-Cu). SELDI ProteinChip mass spectrometry was then performed. SELDI data were analyzed, protein peak clustering and classification were performed, and a supervised classification algorithm was employed to classify the dataset. Multiple protein peaks ranging from 3.3 to 30 kDa were identified between patients with recurrence and those without recurrence, and the expression pattern differences of three proteins were used to generate the discriminating classification tree. The biomarkers were expressed with a high degree of reproducibility. In this early characterization study, melanoma recurrence was predicted with a sensitivity of 72% (18/25) and a specificity of 75% (18/24). This novel pilot study revealed three proteins that accurately identified patients who developed recurrence after curative resection of primary melanoma.

Adult↗

Medulloblastoma and birth date: evaluation of 3 U.S. datasets.

Studies from Norway and Japan have found a higher incidence of medulloblastoma related to births that occur in the fall. The authors sought further evidence concerning this association. For 122 patients in a Duke University database and 90 patients from the Central Cancer Registry of North Carolina, the frequency distribution of birth dates by month was statistically significantly different from the expected North Carolina distribution (p = 0.04 and 0.06). For 75 patients from California Surveillance, Epidemiology, and End Results (SEER) data, the frequency distribution of birth dates by month was marginally different from the expected U.S. distribution (p = 0.14). For 922 patients from national SEER data, the frequency distribution of birth dates by month was not statistically significantly different from the expected U.S. distribution (p = 0.54). Subgroup analysis suggests seasonality of birth dates is most significant for patients aged 5-14 yr diagnosed with medulloblastoma.

Adolescent↗

A clinical evaluation of flurbiprofen LAT and piroxicam gel: a multicentre study in general practice.

A prospective, randomized, multicentre, open, crossover study of the comparative efficacy, tolerability and acceptability of two topical nonsteroidal anti-inflammatory drug (NSAID) therapies, flurbiprofen local-action transcutaneous (LAT) patch (40 mg b.d.) and piroxicam gel (3 cm, 0.5% q.d.s), was conducted in general practice in the UK in 137 men and women with soft-tissue rheumatism of the shoulder or elbow (e.g. epicondylitis, tendinitis, bursitis and adhesive capsulitis). Patients received one therapy for 4 days before crossing over to the other NSAID for a further 4 days, followed by 6 days of their preferred therapy. Clinical assessment of severity of pain, tenderness and overall clinical condition was carried out at baseline and after 4, 8 and 14 days. Patients self-assessed the severity of pain during the day and at night, and also the quality of their sleep during each treatment phase. More patients showed a greater improvement in all of the clinical assessments of efficacy following treatment with flurbiprofen LAT during the crossover phase. There was a statistically significant reduction in the severity of pain, the principal measure of efficacy, in favour of flurbiprofen LAT: 42% of patients showed greater improvement with flurbiprofen LAT compared with 26% who showed a greater improvement with piroxicam gel (p = 0.012; n = 131, intent-to-treat). Eligible dataset (n = 126) analysis revealed statistically significant differences in favour of flurbiprofen LAT in the severity of lesion tenderness (p = 0.03) and the overall change in clinical condition (p = 0.04) compared with baseline status. Superior efficacy for flurbiprofen LAT was also indicated in the patients' assessment at the end of the crossover phase (day 8), at which 69% chose to continue treatment with flurbiprofen LAT compared with only 31% of patients who chose piroxicam gel (n = 126; p < 0.001). There were, in addition, statistically significant differences in favour of flurbiprofen LAT in assessments for night pain (p < 0.001), quality of sleep (p = 0.004) and the patients' overall opinion of treatment (p < 0.001). Both treatments were well tolerated with a low incidence of mainly local adverse events. These results showed that flurbiprofen LAT had a greater efficacy than piroxicam gel, and was also preferred by patients in the treatment of painful soft-tissue rheumatism of the shoulder and elbow.

Administration, Cutaneous↗

eMelanoBase: an online locus-specific variant database for familial melanoma.

A proportion of melanoma-prone individuals in both familial and non-familial contexts has been shown to carry inactivating mutations in either CDKN2A or, rarely, CDK4. CDKN2A is a complex locus that encodes two unrelated proteins from alternately spliced transcripts that are read in different frames. The alpha transcript (exons 1alpha, 2, and 3) produces the p16INK4A cyclin-dependent kinase inhibitor, while the beta transcript (exons 1beta and 2) is translated as p14ARF, a stabilizing factor of p53 levels through binding to MDM2. Mutations in exon 2 can impair both polypeptides and insertions and deletions in exons 1alpha, 1beta, and 2, which can theoretically generate p16INK4A-p14ARF fusion proteins. No online database currently takes into account all the consequences of these genotypes, a situation compounded by some problematic previous annotations of CDKN2A-related sequences and descriptions of their mutations. As an initiative of the international Melanoma Genetics Consortium, we have therefore established a database of germline variants observed in all loci implicated in familial melanoma susceptibility. Such a comprehensive, publicly accessible database is an essential foundation for research on melanoma susceptibility and its clinical application. Our database serves two types of data as defined by HUGO. The core dataset includes the nucleotide variants on the genomic and transcript levels, amino acid variants, and citation. The ancillary dataset includes keyword description of events at the transcription and translation levels and epidemiological data. The application that handles users' queries was designed in the model-view-controller architecture and was implemented in Java. The object-relational database schema was deduced using functional dependency analysis. We hereby present our first functional prototype of eMelanoBase. The service is accessible via the URL www.wmi.usyd.edu.au:8080/melanoma.html.

Computer Security↗

Transparent image access in a distributed picture archiving and communications system: the Master Database broker.

A distributed design is the most cost-effective system for small-to medium-scale picture archiving and communications systems (PACS) implementations. However, the design presents an interesting challenge to developers and implementers: to make stored image data, distributed throughout the PACS network, appear to be centralized with a single access point for users. A key component for the distributed system is a central or master database, containing all the studies that have been scanned into the PACS. Each study includes a list of one or more locations for that particular dataset so that applications can easily find it. Non-Digital Imaging and Communications in Medicine (DICOM) clients, such as our worldwide web (WWW)-based PACS browser, query the master database directly to find the images, then jump to the most appropriate location via a distributed web-based viewing system. The Master Database Broker provides DICOM clients with the same functionality by translating DICOM queries to master database searches and distributing retrieval requests transparently to the appropriate source. The Broker also acts as a storage service class provider, allowing users to store selected image subsets and reformatted images with the original study, without having to know on which server the original data are stored.

CD-ROM↗

Visual detection of spatial contrast patterns: evaluation of five simple models.

The ModelFest Phase One dataset is a collection of luminance contrast thresholds for 43 two-dimensional monochromatic spatial patterns confined to an area of approximately two by two degrees. These data were collected by a collaboration among twelve laboratories, and were designed to provide a common database for calibration and testing of spatial vision models. Here I report fits of the ModelFest data with five models: Peak Contrast, Contrast Energy, Generalized Energy, a Gabor Channels model, and a Discrete Cosine Transform model. The Gabor Channels model provides the best fit, though the other, simpler models, with the exception of Peak Contrast, provide remarkably good fits as well. Though there are clear individual differences, regularities in the data suggest the possibility of constructing a standard observer for spatial vision.

Contrast Sensitivity↗

Outcome signature genes in breast cancer: is there a unique set?

MOTIVATION: Predicting the metastatic potential of primary malignant tissues has direct bearing on the choice of therapy. Several microarray studies yielded gene sets whose expression profiles successfully predicted survival. Nevertheless, the overlap between these gene sets is almost zero. Such small overlaps were observed also in other complex diseases, and the variables that could account for the differences had evoked a wide interest. One of the main open questions in this context is whether the disparity can be attributed only to trivial reasons such as different technologies, different patients and different types of analyses. RESULTS: To answer this question, we concentrated on a single breast cancer dataset, and analyzed it by a single method, the one which was used by van't Veer et al. to produce a set of outcome-predictive genes. We showed that, in fact, the resulting set of genes is not unique; it is strongly influenced by the subset of patients used for gene selection. Many equally predictive lists could have been produced from the same analysis. Three main properties of the data explain this sensitivity: (1) many genes are correlated with survival; (2) the differences between these correlations are small; (3) the correlations fluctuate strongly when measured over different subsets of patients. A possible biological explanation for these properties is discussed. CONTACT: eytan.domany@weizmann.ac.il SUPPLEMENTARY INFORMATION: http://www.weizmann.ac.il/physics/complex/compphys/downloads/liate/

Biomarkers, Tumor↗

Peak flow monitoring in clinical practice and clinical asthma trials.

PURPOSE OF REVIEW: This review summarizes recent reports on peak expiratory flow (PEF) monitoring in clinical asthma trials and clinical practice. RECENT FINDINGS: In clinical trials, summary measures such as average morning PEF provide only a fraction of the available information about asthma control and treatment response. New statistical models should improve the yield from PEF datasets. Improved criteria are needed for the diagnosis of exacerbations, and these may be developed from quality-control analysis of existing datasets. In clinical practice we must reduce the burden of monitoring and increase the ease of interpretation of PEF data. Electronic monitoring, with short message service or Internet communication, may assist with both. There is a need for standardized user-friendly PEF charts and simple statistically appropriate interpretative tools, which will facilitate the development of clinical algorithms and individualized written action plans. Normal values for diurnal variability should be updated to reflect twice daily monitoring. SUMMARY: Current use of PEF data is limited by the burden of monitoring and the continuing use of interpretative tools that were originally developed for their practical feasibility rather than their clinical validity. Both of these problems may be improved by giving attention to methods for recording, displaying and analysing PEF data.

Asthma↗

On global-local artificial neural networks for function approximation.

We present a hybrid radial basis function (RBF) sigmoid neural network with a three-step training algorithm that utilizes both global search and gradient descent training. The algorithm used is intended to identify global features of an input-output relationship before adding local detail to the approximating function. It aims to achieve efficient function approximation through the separate identification of aspects of a relationship that are expressed universally from those that vary only within particular regions of the input space. We test the effectiveness of our method using five regression tasks; four use synthetic datasets while the last problem uses real-world data on the wave overtopping of seawalls. It is shown that the hybrid architecture is often superior to architectures containing neurons of a single type in several ways: lower mean square errors are often achievable using fewer hidden neurons and with less need for regularization. Our global-local artificial neural network (GL-ANN) is also seen to compare favorably with both perceptron radial basis net and regression tree derived RBFs. A number of issues concerning the training of GL-ANNs are discussed: the use of regularization, the inclusion of a gradient descent optimization step, the choice of RBF spreads, model selection, and the development of appropriate stopping criteria.

Algorithms↗

[Benign and malignant lesions in the region of the inner ear and cerebellopontine angle].

Tumorous lesions in the region of the inner ear and cerebellopontine angle are very rare and can be classified into benign and malignant disease forms. This contribution presents and explains the CT and MRI characteristics of these tumors.High-resolution computed tomography (HRCT) in the axial projection is applied for evaluation in the high-resolution bone window. The coronary slices can be reconstructed from the axial datasets or in individual cases examined in the coronary plane.HRCT excellently demonstrates osseous lesions and in individual cases - e.g., exostoses - it can simply suffice to perform HRCT of the temporal bone, while HRCT is also excellent for detecting osseous lesions to determine whether the tumor is benign or malignant.MRI, on the other hand, excellently shows the extent of tumor spread because of its superb soft tissue contrast. Consequently, HRCT and MRI images of the inner ear and cerebellopontine angle provide meaningful information for visualization and classification of tumorous lesions. The two methods should not be considered as competing but rather as complementary and among other aspects exert considerable influence on the therapeutic approach.

Cerebellar Neoplasms↗

Using GO-PseAA predictor to identify membrane proteins and their types.

Cell membranes are crucial to the life of a cell. Although the basic structure of biological membrane is provided by the lipid bilayer, most of the specific functions are carried out by membrane proteins. Knowledge of membrane protein type often offers important clues toward determining the function of an uncharacterized protein. Therefore, predicting the type of a membrane protein from its primary sequence, or even just identifying whether the uncharacterized protein belongs to a membrane protein or not, is an important and challenging problem in bioinformatics and proteomics. To deal with these problems, the GO-PseAA predictor is introduced that is operated in a hybridization space by combining the gene ontology and pseudo amino acid composition. Meanwhile, to test the prediction quality, a dataset was constructed that contains 6476 non-membrane proteins and 5122 membrane proteins classified into five different types. To avoid redundancy and bias, none of the proteins included has > or = 40% sequence identity to any other. It has been observed that the overall success rate by the jackknife cross-validation test in identifying non-membrane proteins and membrane proteins was 94.76%, and that in identifying the five membrane protein types was 95.84%. The high success rates suggest that the GO-PseAA predictor can catch the core feature of the statistical samples concerned and may become an automated high throughput toll in molecular and cell biology.

Algorithms↗