Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

Prediction of protein structural class by amino acid and polypeptide composition.

A new approach of predicting structural classes of protein domain sequences is presented in this paper. Besides the amino acid composition, the composition of several dipeptides, tripeptides, tetrapeptides, pentapeptides and hexapeptides are taken into account based on the stepwise discriminant analysis. The result of jackknife test shows that this new approach can lead to higher predictive sensitivity and specificity for reduced sequence similarity datasets. Considering the dataset PDB40-B constructed by Brenner and colleagues, 75.2% protein domain sequences are correctly assigned in the jackknife test for the four structural classes: all-alpha, all-beta, alpha/beta and alpha + beta, which is improved by 19.4% in jackknife test and 25.5% in resubstitution test, in contrast with the component-coupled algorithm using amino acid composition alone (AAC approach) for the same dataset. In the cross-validation test with dataset PDB40-J constructed by Park and colleagues, more than 80% predictive accuracy is obtained. Furthermore, for the dataset constructed by Chou and Maggiona, the accuracy of 100% and 99.7% can be easily achieved, respectively, in the resubstitution test and in the jackknife test merely taking the composition of dipeptides into account. Therefore, this new method provides an effective tool to extract valuable information from protein sequences, which can be used for the systematic analysis of small or medium size protein sequences. The computer programs used in this paper are available on request.

Algorithms↗

Prospective development of a cardiac surgical registry.

Interest has recently been expressed in developing an Australian adult cardiac surgical registry. Complete national registries of adult cardiac surgery have already been established in many European countries, the USA, Canada and elsewhere. Participating centres contributing to a national registry benefit by being able to benchmark themselves against norms for their particular country. A risk-adjusted database can help surgeons advise their patients of the chances of a good operative outcome. For a surgeon or a surgical unit, the only way to obtain a relevant risk model is to use their own data, and data from units in their particular country. It is also useful to have comparative data from other national registries to compare one's own country with international benchmarks. Since 1996, the European Cardiac Surgical Registry (ECSUR) has put considerable effort into producing unified datasets, harmonised with each other for worldwide use. In 1997, ECSUR launched a minimum cardiac surgical dataset. The worldwide launch of the full international adult cardiac surgical dataset is scheduled for July/August 2000. This dataset would be highly useful for application in Australia. The ECSUR organisation has the capability to analyse data from other countries and could perform this for Australia if requested. However, a better approach would be a national centre in Australia. Funding for national registries around the world has been obtained from Ministries of Health, participating surgical centres, and surgical software vendors. If an Australian national registry is indeed established it will find a ready-made, highly appropriate international cardiac surgical dataset sponsored by ECSUR and the Society of Thoracic Surgeons waiting for adoption by Australia.

Journal Article↗

Intrinsic disorder in the Protein Data Bank.

The Protein Data Bank (PDB) is the preeminent source of protein structural information. PDB contains over 32,500 experimentally determined 3-D structures solved using X-ray crystallography or nuclear magnetic resonance spectroscopy. Intrinsically disordered regions fail to form a fixed 3-D structure under physiological conditions. In this study, we compare the amino-acid sequences of proteins whose structures are determined by X-ray crystallography with the corresponding sequences from the Swiss-Prot database. The analyzed dataset includes 16,370 structures, which represent 18,101 PDB chains and 5,434 different proteins from 910 different organisms (2,793 eukaryotic, 2,109 bacterial, 288 viral, and 244 archaeal). In this dataset, on average, each Swiss-Prot protein is represented by 7 PDB chains with 76% of the crystallized regions being represented by more than one structure. Intriguingly, the complete sequences of only approximately 7% of proteins are observed in the corresponding PDB structures, and only approximately 25% of the total dataset have >95% of their lengths observed in the corresponding PDB structures. This suggests that the vast majority of PDB proteins is shorter than their corresponding Swiss-Prot sequences and/or contain numerous residues, which are not observed in maps of electron density. To determine the prevalence of disordered regions in PDB, the residues in the Swiss-Prot sequences were grouped into four general categories, "Observed" (which correspond to structured regions), "Not observed" (regions with missing electron density, potentially disordered), "Uncharacterized," and "Ambiguous," depending on their appearance in the corresponding PDB entries. This non-redundant set of residues can be viewed as a 'fragment' or empirical domain database that contains a set of experimentally determined structured regions or domains and a set of experimentally verified disordered regions or domains. We studied the propensities and properties of residues in these four categories and analyzed their relations to the predictions of disorder using several algorithms. "Non-observed," "Ambiguous," and "Uncharacterized" regions were shown to possess the amino acid compositional biases typical of intrinsically disordered proteins. The application of four different disorder predictors (PONDR(R) VL-XT, VL3-BA, VSL1P, and IUPred) revealed that the vast majority of residues in the "Observed" dataset are ordered, and that the "Not observed" regions are mostly disordered. The "Uncharacterized" regions possess some tendency toward order, whereas the predictions for the short "Ambiguous" regions are really ambiguous. Long "Ambiguous" regions (>70 amino acid residues) are mostly predicted to be ordered, suggesting that they are likely to be "wobbly" domains. Overall, we showed that completely ordered proteins are not highly abundant in PDB and many PDB sequences have disordered regions. In fact, in the analyzed dataset approximately 10% of the PDB proteins contain regions of consecutive missing or ambiguous residues longer than 30 amino-acids and approximately 40% of the proteins possess short regions (> or =10 and < 30 amino-acid long) of missing and ambiguous residues.

Algorithms↗

Four-dimensional fetal echocardiography with spatiotemporal image correlation (STIC): a systematic study of standard cardiac views assessed by different observers.

OBJECTIVE: To test the agreement between observers and reproducibility of a technique to display standard cardiac views of the left and right ventricular outflow tracts from four-dimensional volume datasets acquired with Spatiotemporal Image Correlation (STIC). METHODS: A technique was developed to obtain dynamic multiplanar images of the left ventricular outflow tract (LVOT) and right ventricular outflow tract (RVOT) from volume datasets acquired with STIC. Volume datasets were acquired from fetuses with normal cardiac anatomy. Twenty volume datasets of satisfactory quality were pre-selected by one investigator. The data was randomly assigned for a blinded review by two independent observers with previous experience in fetal echocardiography. Only one volume dataset was used for each fetus. After a training session, the observers obtained standardized cardiac views of the LVOT and RVOT, which were scored on a scale of 1 to 5, based on diagnostic value and image quality (1=unacceptable, 2=marginal, 3=acceptable, 4=good, and 5=excellent). Median scores and interquartile range, as well as inter- and intraobserver agreement were calculated for each view. RESULTS: The mean menstrual age at the time of volume acquisition was 25.5+/-4.5 weeks. Median scores (interquartile range) for LVOT images, obtained by the first and second observers, were 3.5 (2.25-5.00) and 4 (3.00-5.00), respectively. The median scores (interquartile range) for RVOT images obtained by the first and second observers were 3 (3.00-5.00) and 3 (2.00-4.00), respectively. The interobserver intraclass correlation coefficient for the LVOT was 0.693 (95% CI 0.380-0.822), and 0.696 (95% CI 0.382-0.866) for the RVOT. For the intraobserver agreement analysis, observer 1 gave higher scores to the LVOT the second time the volumes were analyzed [LVOT: 3.50 (2.25-5.00) vs. 5.00 (4.00-5.00, p=0.008)]. CONCLUSION: STIC can be reproducibly used to evaluate fetal cardiac outflow tracts by independent examiners. Slightly better image quality rating scores during the intraobserver variability trial suggests the presence of a learning curve for the manipulation and analysis of volume data obtained by STIC.

Double-Blind Method↗

Evaluating methods for classifying expression data.

An attractive application of expression technologies is to predict drug efficacy or safety using expression data of biomarkers. To evaluate the performance of various classification methods for building predictive models, we applied these methods on six expression datasets. These datasets were from studies using microarray technologies and had either two or more classes. From each of the original datasets, two subsets were generated to simulate two scenarios in biomarker applications. First, a 50-gene subset was used to simulate a candidate gene approach when it might not be practical to measure a large number of genes/biomarkers. Next, a 2000-gene subset was used to simulate a whole genome approach. We evaluated the relative performance of several classification methods by using leave-one-out cross-validation and bootstrap cross-validation. Although all methods perform well in both subsets for a relative easy dataset with two classes, differences in performance do exist among methods for other datasets. Overall, partial least squares discriminant analysis (PLS-DA) and support vector machines (SVM) outperform all other methods. We suggest a practical approach to take advantage of multiple methods in biomarker applications.

Algorithms↗

An enhanced block matching algorithm for fast elastic registration in adaptive radiotherapy.

Image registration has many medical applications in diagnosis, therapy planning and therapy. Especially for time-adaptive radiotherapy, an efficient and accurate elastic registration of images acquired for treatment planning, and at the time of the actual treatment, is highly desirable. Therefore, we developed a fully automatic and fast block matching algorithm which identifies a set of anatomical landmarks in a 3D CT dataset and relocates them in another CT dataset by maximization of local correlation coefficients in the frequency domain. To transform the complete dataset, a smooth interpolation between the landmarks is calculated by modified thin-plate splines with local impact. The concept of the algorithm allows separate processing of image discontinuities like temporally changing air cavities in the intestinal track or rectum. The result is a fully transformed 3D planning dataset (planning CT as well as delineations of tumour and organs at risk) to a verification CT, allowing evaluation and, if necessary, changes of the treatment plan based on the current patient anatomy without time-consuming manual re-contouring. Typically the total calculation time is less than 5 min, which allows the use of the registration tool between acquiring the verification images and delivering the dose fraction for online corrections. We present verifications of the algorithm for five different patient datasets with different tumour locations (prostate, paraspinal and head-and-neck) by comparing the results with manually selected landmarks, visual assessment and consistency testing. It turns out that the mean error of the registration is better than the voxel resolution (2 x 2 x 3 mm(3)). In conclusion, we present an algorithm for fully automatic elastic image registration that is precise and fast enough for online corrections in an adaptive fractionated radiation treatment course.

Algorithms↗

Probabilistic disease classification of expression-dependent proteomic data from mass spectrometry of human serum.

We have developed an algorithm called Q5 for probabilistic classification of healthy versus disease whole serum samples using mass spectrometry. The algorithm employs principal components analysis (PCA) followed by linear discriminant analysis (LDA) on whole spectrum surface-enhanced laser desorption/ionization time of flight (SELDI-TOF) mass spectrometry (MS) data and is demonstrated on four real datasets from complete, complex SELDI spectra of human blood serum. Q5 is a closed-form, exact solution to the problem of classification of complete mass spectra of a complex protein mixture. Q5 employs a probabilistic classification algorithm built upon a dimension-reduced linear discriminant analysis. Our solution is computationally efficient; it is noniterative and computes the optimal linear discriminant using closed-form equations. The optimal discriminant is computed and verified for datasets of complete, complex SELDI spectra of human blood serum. Replicate experiments of different training/testing splits of each dataset are employed to verify robustness of the algorithm. The probabilistic classification method achieves excellent performance. We achieve sensitivity, specificity, and positive predictive values above 97% on three ovarian cancer datasets and one prostate cancer dataset. The Q5 method outperforms previous full-spectrum complex sample spectral classification techniques and can provide clues as to the molecular identities of differentially expressed proteins and peptides.

Algorithms↗

Default values for assessment of potential dermal exposure of the hands to industrial chemicals in the scope of regulatory risk assessments.

Dermal exposure needs to be addressed in regulatory risk assessment of chemicals. The models used so far are based on very limited data. The EU project RISKOFDERM has gathered a large number of new measurements on dermal exposure to industrial chemicals in various work situations, together with information on possible determinants of exposure. These data and information, together with some non-RISKOFDERM data were used to derive default values for potential dermal exposure of the hands for so-called 'TGD exposure scenarios'. TGD exposure scenarios have similar values for some very important determinant(s) of dermal exposure, such as amount of substance used. They form narrower bands within the so-called 'RISKOFDERM scenarios', which cluster exposure situations according to the same purpose of use of the products. The RISKOFDERM scenarios in turn are narrower bands within the so-called Dermal Exposure Operation units (DEO units) that were defined in the RISKOFDERM project to cluster situations with similar exposure processes and exposure routes. Default values for both reasonable worst case situations and typical situations were derived, both for single datasets and, where possible, for combined datasets that fit the same TGD exposure scenario. The following reasonable worst case potential hand exposures were derived from combined datasets: (i) loading and filling of large containers (or mixers) with large amounts (many litres) of liquids: 11,500 mg per scenario (14 mg cm(-2) per scenario with surface of the hands assumed to be 820 cm(2)); (ii) careful mixing of small quantities (tens of grams in <1l): 4.1 mg per scenario (0.005 mg cm(-2) per scenario); (iii) spreading of (viscous) liquids with a comb on a large surface area: 130 mg per scenario (0.16 mg cm(-2) per scenario); (iv) brushing and rolling of (relatively viscous) liquid products on surfaces: 6500 mg per scenario (8 mg cm(-2) per scenario) and (v) spraying large amounts of liquids (paints, cleaning products) on large areas: 12,000 mg per scenario (14 mg cm(-2) per scenario). These default values are considered useful for estimating exposure for similar substances in similar situations with low uncertainty. Several other default values based on single datasets can also be used, but lead to estimates with a higher uncertainty, due to their more limited basis. Sufficient analogy in all described parameters of the scenario, including duration, is needed to enable proper use of the default values. The default values lead to similar estimates as the RISKOFDERM dermal exposure model that was based on the same datasets, but uses very different parameters. Both approaches are preferred over older general models, such as EASE, that are not based on data from actual dermal exposure situations.

Environmental Monitoring↗

GeneRAGE: a robust algorithm for sequence clustering and domain detection.

MOTIVATION: Efficient, accurate and automatic clustering of large protein sequence datasets, such as complete proteomes, into families, according to sequence similarity. Detection and correction of false positive and negative relationships with subsequent detection and resolution of multi-domain proteins. RESULTS: A new algorithm for the automatic clustering of protein sequence datasets has been developed. This algorithm represents all similarity relationships within the dataset in a binary matrix. Removal of false positives is achieved through subsequent symmetrification of the matrix using a Smith-Waterman dynamic programming alignment algorithm. Detection of multi-domain protein families and further false positive relationships within the symmetrical matrix is achieved through iterative processing of matrix elements with successive rounds of Smith-Waterman dynamic programming alignments. Recursive single-linkage clustering of the corrected matrix allows efficient and accurate family representation for each protein in the dataset. Initial clusters containing multi-domain families, are split into their constituent clusters using the information obtained by the multi-domain detection step. This algorithm can hence quickly and accurately cluster large protein datasets into families. Problems due to the presence of multi-domain proteins are minimized, allowing more precise clustering information to be obtained automatically. AVAILABILITY: GeneRAGE (version 1.0) executable binaries for most platforms may be obtained from the authors on request. The system is available to academic users free of charge under license.

Algorithms↗

CeLLTra: aligning cell names with gene expression via a pathway-informed transformer.

MOTIVATION: Single-cell RNA sequencing (scRNA-Seq) technology enables detailed exploration of gene expression at the individual cell level, crucial for annotating cell types and understanding cellular diversity. Traditional methods for cell type annotation often rely on marker genes and manual labeling, posing challenges due to low data quality and incomplete reference datasets. RESULTS: We developed CeLLTra, a novel contrastive learning framework that leverages a Transformer-based model integrating biological pathway information to group genes into super tokens, effectively capturing comprehensive gene expression from scRNA-Seq data. By combining this pathway-informed Transformer with a pretrained domain-specific language model, CeLLTra accurately aligns cell-type annotations with gene expression profiles. Evaluations on a large-scale human scRNA-Seq dataset showed that CeLLTra significantly outperformed state-of-the-art methods in supervised and zero-shot cell-type prediction. Additionally, CeLLTra generalized well to external datasets, improving clustering performance and enabling better characterization of cancerous cell states in tumor-infiltrating myeloid cells from non-small cell lung cancer patients. AVAILABILITY AND IMPLEMENTATION: CeLLTra is freely available on GitHub (https://github.com/WJZheng-group/CeLLTra) and Zenodo (https://doi.org/10.5281/zenodo.17666735). The datasets underlying this article are the following: GSE201333 and GSE127465. All these datasets are publicly available and can be freely accessed on the Gene Expression Omnibus repository.

Humans↗

Evaluation of epistasis detection methods for quantitative phenotypes.

MOTIVATION: Epistasis, or genetic interaction, plays a crucial role in shaping complex traits and has been increasingly recognized for its widespread influence in genetic architectures. While epistasis detection has been extensively evaluated in case-control studies, its performance with quantitative phenotypes remains comparatively understudied. RESULTS: We identified and evaluated six epistasis detection methods applicable to quantitative trait analysis: EpiSNP, Matrix Epistasis, MIDESP, PLINK Epistasis, QMDR, and REMMA. Using the EpiGEN simulator, we generated synthetic datasets modeling four classes of pairwise SNP interactions-dominant, multiplicative, recessive, and XOR. We also assessed BOOST and MDR algorithms using discretized (case-control) versions of the same datasets. Performance varied notably by interaction type: REMMA achieved the highest overall detection rate (55%), particularly excelling with dominant interactions (100%). MDR excelled with multiplicative (57%) and XOR (69%) interactions. Meanwhile, EpiSNP attained the best performance for recessive interactions (67%). All methods except BOOST produced F1 scores below 0.05 for most interaction types. We further evaluated the methods using a real-world dataset. When applied to the Adolescent Brain Cognitive Development dataset to analyse the externalizing behavior phenotype, both PLINK Epistasis and PLINK BOOST identified SNPs within the DRD2 and DRD4 genes, consistent with previously reported genetic associations. Given the variability in tool performance across interaction types, no single method provides optimal detection across all scenarios. Leveraging multiple detection algorithms may therefore yield more comprehensive insights into epistatic effects in quantitative trait analyses. AVAILABILITY AND IMPLEMENTATION: All relevant code and simulated datasets can be found at github.com/staslist/Epistasis_Review repository.

Epistasis, Genetic↗

Gene co-expression network topology provides a framework for molecular characterization of cellular state.

MOTIVATION: Gene expression data have become an instrumental resource in describing the molecular state associated with various cellular phenotypes and responses to environmental perturbations. The utility of expression profiling has been demonstrated in partitioning clinical states, predicting the class of unknown samples and in assigning putative functional roles to previously uncharacterized genes based on profile similarity. However, gene expression profiling has had only limited success in identifying therapeutic targets. This is partly due to the fact that current methods based on fold-change focus only on single genes in isolation, and thus cannot convey causal information. In this paper, we present a technique for analysis of expression data in a graph-theoretic framework that relies on associations between genes. We describe the global organization of these networks and biological correlates of their structure. We go on to present a novel technique for the molecular characterization of disparate cellular states that adds a new dimension to the fold-based methods and conclude with an example application to a human medulloblastoma dataset. RESULTS: We have shown that expression networks generated from large model-organism expression datasets are scale-free and that the average clustering coefficient of these networks is several orders of magnitude higher than would be expected for similarly sized scale-free networks, suggesting an inherent hierarchical modularity similar to that previously identified in other biological networks. Furthermore, we have shown that these properties are robust with respect to the parameters of network construction. We have demonstrated an enrichment of genes having lethal knockout phenotypes in the high-degree (i.e. hub) nodes in networks generated from aggregate condition datasets; using process-focused Saccharomyces cerivisiae datasets we have demonstrated additional high-degree enrichments of condition-specific genes encoding proteins known to be involved in or important for the processes interrogated by the microarrays. These results demonstrate the utility of network analysis applied to expression data in identifying genes that are regulated in a state-specific manner. We concluded by showing that a sample application to a human clinical dataset prominently identified a known therapeutic target. AVAILABILITY: Software implementing the methods for network generation presented in this paper is available for academic use by request from the authors in the form of compiled linux binary executables.

Algorithms↗

Regulatory motif finding by logic regression.

MOTIVATION: Multiple transcription factors coordinately control transcriptional regulation of genes in eukaryotes. Although many computational methods consider the identification of individual transcription factor binding sites (TFBSs), very few focus on the interactions between these sites. We consider finding TFBSs and their context specific interactions using microarray gene expression data. We devise a hybrid approach called LogicMotif composed of a TFBS identification method combined with the new regression methodology logic regression. LogicMotif has two steps: First, potential binding sites are identified from transcription control regions of genes of interest. Various available methods can be used in this step when the genes of interest can be divided into groups such as up-and downregulated. For this step, we also develop a simple univariate regression and extension method MFURE to extract candidate TFBSs from a large number of genes in the availability of microarray gene expression data. MFURE provides an alternative method for this step when partitioning of the genes into disjoint groups is not preferred. This first step aims to identify individual sites within gene groups of interest or sites that are correlated with the gene expression outcome. In the second step, logic regression is used to build a predictive model of outcome of interest (either gene expression or up- and down-regulation) using these potential sites. This 2-fold approach creates a rich diverse set of potential binding sites in the first step and builds regression or classification models in the second step using logic regression that is particularly good at identifying complex interactions. RESULTS: LogicMotif is applied to two publicly available datasets. A genome-wide gene expression data set of Saccharomyces cerevisiae is used for validation. The regression models obtained are interpretable and the biological implications are in agreement with the known resuts. This analysis suggests that LogicMotif provides biologically more reasonable regression models than previous analysis of this dataset with standard linear regression methods. Another dataset of S.cerevisiae illustrates the use of LogicMotif in classification questions by building a model that discriminates between up- and down-regulated genes in iron copper deficiency. LogicMotif identifies an inductive and two repressor motifs in this dataset. The inductive motif matches the binding site of the transcription factor Aft1p that has a key role in regulation of the uptake process. One of the novel repressor sites is highly present in transcription control regions of FeS genes. This site could represent a TFBS for an unknown transcription factor involved in repression of genes encoding FeS proteins in iron deficiency. We establish the robustness of the method to the type of outcome variable used by considering both continuous and binary outcome variables for this dataset. Our results indicate that logic regression used in combination with cluster/group operating binding site identification methods or with our proposed method MFURE is a powerful and flexible alternative to linear regression based motif finding methods. AVAILABILITY: Source code for logic regression is freely available as a package of the R programming language by Ruczinski et al. (2003) and can be downloaded at http://bear.fhcrc.org/~ingor/logic/download/download.html an R package for MFURE is available at http://www.stat.berkeley.edu/~sunduz/software.html

Algorithms↗

Analysis of domain correlations in yeast protein complexes.

MOTIVATION: A growing body of research has concentrated on the identification and definition of conserved sequence motifs. It is widely recognized that these conserved sequence and structural units often mediate protein functions and interactions. The continuing advancements in high-throughput experiments necessitate the development of computational methods to critically assess the results. In this work, we analyzed high-throughput protein complexes using the domain composition of their protein constituents. Domains that mediate similar or related functions may consistently co-occur in protein complexes. RESULTS: We analyzed Saccharomyces cerevisiae protein complexes from curated and high-throughput experimental datasets to identify statistically significant functional associations between domains. The resulting correlations are represented as domain networks that form the basis of comparison between the datasets, as well as to binary protein interactions. The results show that the curated datasets produce domain networks that map to known biological assemblies, such as ribosome, RNA polymerase, proteasome regulators, transcription initiation and histones. Furthermore, many of these domain correlations were also found in binary protein interactions. In contrast, the high-throughput datasets contain one large network of domain associations. High connectivity of RNA processing and binding domains in the high-throughput datasets reflects the abundance of RNA binding proteins in yeast, in agreement with a previous report that identified a nucleolar protein cluster, possibly mediated by rRNA, from these complexes. AVAILABILITY: The software is available upon request from the authors and is dependent on the NCBI C++ toolkit.

Amino Acid Motifs↗

A probabilistic model for mining implicit 'chemical compound-gene' relations from literature.

MOTIVATION: The importance of chemical compounds has been emphasized more in molecular biology, and 'chemical genomics' has attracted a great deal of attention in recent years. Thus an important issue in current molecular biology is to identify biological-related chemical compounds (more specifically, drugs) and genes. Co-occurrence of biological entities in the literature is a simple, comprehensive and popular technique to find the association of these entities. Our focus is to mine implicit 'chemical compound and gene' relations from the co-occurrence in the literature. RESULTS: We propose a probabilistic model, called the mixture aspect model (MAM), and an algorithm for estimating its parameters to efficiently handle different types of co-occurrence datasets at once. We examined the performance of our approach not only by a cross-validation using the data generated from the MEDLINE records but also by a test using an independent human-curated dataset of the relationships between chemical compounds and genes in the ChEBI database. We performed experimentation on three different types of co-occurrence datasets (i.e. compound-gene, gene-gene and compound-compound co-occurrences) in both cases. Experimental results have shown that MAM trained by all datasets outperformed any simple model trained by other combinations of datasets with the difference being statistically significant in all cases. In particular, we found that incorporating compound-compound co-occurrences is the most effective in improving the predictive performance. We finally computed the likelihoods of all unknown compound-gene (more specifically, drug-gene) pairs using our approach and selected the top 20 pairs according to the likelihoods. We validated them from biological, medical and pharmaceutical viewpoints.

Artificial Intelligence↗

Support vector machine learning from heterogeneous data: an empirical analysis using protein sequence and structure.

MOTIVATION: Drawing inferences from large, heterogeneous sets of biological data requires a theoretical framework that is capable of representing, e.g. DNA and protein sequences, protein structures, microarray expression data, various types of interaction networks, etc. Recently, a class of algorithms known as kernel methods has emerged as a powerful framework for combining diverse types of data. The support vector machine (SVM) algorithm is the most popular kernel method, due to its theoretical underpinnings and strong empirical performance on a wide variety of classification tasks. Furthermore, several recently described extensions allow the SVM to assign relative weights to various datasets, depending upon their utilities in performing a given classification task. RESULTS: In this work, we empirically investigate the performance of the SVM on the task of inferring gene functional annotations from a combination of protein sequence and structure data. Our results suggest that the SVM is quite robust to noise in the input datasets. Consequently, in the presence of only two types of data, an SVM trained from an unweighted combination of datasets performs as well or better than a more sophisticated algorithm that assigns weights to individual data types. Indeed, for this simple case, we can demonstrate empirically that no solution is significantly better than the naive, unweighted average of the two datasets. On the other hand, when multiple noisy datasets are included in the experiment, then the naive approach fares worse than the weighted approach. Our results suggest that for many applications, a naive unweighted sum of kernels may be sufficient. AVAILABILITY: http://noble.gs.washington.edu/proj/seqstruct

Algorithms↗

Maternal role in type 2 diabetes mellitus: indirect evidence for a mitochondrial inheritance.

BACKGROUND: The role of mitochondrial inheritance in type 2 diabetes mellitus has received much attention recently. In this study, three existing datasets in Taiwan are analysed to examine this theory. METHODS: Two of the datasets were community surveys and one a hospital case series. Subjects who had information regarding their paternal or maternal diabetic status were selected for the present study. In the first dataset, 745 subjects had information about paternal diabetic status and 765 had information about maternal diabetic status. In the second dataset, 255 and 267 subjects had the paternal and maternal information, respectively. In the third, a total of 3625 subjects had information about both their paternal and maternal diabetic status. Diabetic status of the study subjects was determined by fasting plasma glucose levels and/or oral glucose tolerance test; their parental diabetic status was collected by interview. RESULTS: The three datasets consistently demonstrated a significantly elevated odds ratio (OR) for reporting maternal diabetes (OR = 2.64, 95% confidence interval: 1.12-5.71) in diabetic patients as compared to non-diabetic subjects. The reporting of paternal diabetes, however, was not significantly different between diabetics and non-diabetics. In addition, the OR for reporting paternal diabetes were not significantly different between groups of different age-at-onset. With respect to maternal history, the OR increased significantly when age-at-onset was younger (test-for-trend P < 0.001). CONCLUSIONS: Our findings support mitochondrial inheritance of type 2 diabetes mellitus in the human population. The age-modifying effect was also in accordance with the mitochondrial oxidative phosphorylation paradigm for degenerative diseases.

Adult↗

Signal in noise: evaluating reported reproducibility of serum proteomic tests for ovarian cancer.

Proteomic profiling of serum initially appeared to be dramatically effective for diagnosis of early-stage ovarian cancer, but these results have proven difficult to reproduce. A recent publication reported good classification in one dataset using results from training on a much earlier dataset, but the authors have since reported that they did not perform the analysis as described. We examined the reproducibility of the proteomic patterns across datasets in more detail. Our analysis reveals that the pattern that enabled successful classification is biologically implausible and that the method, properly applied, does not classify the data accurately. We show that the method used in previously published studies does not establish reproducibility and performs no better than chance for classifying the second dataset, in part because the second dataset is easy to classify correctly. We conclude that the reproducibility of the proteomic profiling approach has yet to be established.

Female↗