Search PubMed⌕ Search

Biomedical subjects

Axel Facius

Publications and source records attributed to Axel Facius.

5 recordsLinked to original sources

Super paramagnetic clustering of protein sequences.

BACKGROUND: Detection of sequence homologues represents a challenging task that is important for the discovery of protein families and the reliable application of automatic annotation methods. The presence of domains in protein families of diverse function, inhomogeneity and different sizes of protein families create considerable difficulties for the application of published clustering methods. RESULTS: Our work analyses the Super Paramagnetic Clustering (SPC) and its extension, global SPC (gSPC) algorithm. These algorithms cluster input data based on a method that is analogous to the treatment of an inhomogeneous ferromagnet in physics. For the SwissProt and SCOP databases we show that the gSPC improves the specificity and sensitivity of clustering over the original SPC and Markov Cluster algorithm (TRIBE-MCL) up to 30%. The three algorithms provided similar results for the MIPS FunCat 1.3 annotation of four bacterial genomes, Bacillus subtilis, Helicobacter pylori, Listeria innocua and Listeria monocytogenes. However, the gSPC covered about 12% more sequences compared to the other methods. The SPC algorithm was programmed in house using C++ and it is available at http://mips.gsf.de/proj/spc. The FunCat annotation is available at http://mips.gsf.de. CONCLUSION: The gSPC calculated to a higher accuracy or covered a larger number of sequences than the TRIBE-MCL algorithm. Thus it is a useful approach for automatic detection of protein families and unsupervised annotation of full genomes.

Algorithms↗

PRIME: a graphical interface for integrating genomic/proteomic databases.

Data mining, finding and integration of information about proteins of interest, is an essential component in modern biological and biomedical research. Even when focusing on a single organism and only on a small number of proteins, there are often dozens fo data sources containing relevant information. We are developing PRIME, a protein information environment, to serve as a virtual central database which integrates distributed heterogeneous information about proteins (linked by common identifier). PRIME has powerful capabilities to visualize all kinds of protein annotation in specialized views. These views can be displayed side by side at the same time and can be synchronized in order to show simultaneously different aspects of identical proteins. These features allow a quick and comprehensive overview of properties of single proteins or protein sets.

Computational Biology↗

Gene selection from microarray data for cancer classification--a machine learning approach.

A DNA microarray can track the expression levels of thousands of genes simultaneously. Previous research has demonstrated that this technology can be useful in the classification of cancers. Cancer microarray data normally contains a small number of samples which have a large number of gene expression levels as features. To select relevant genes involved in different types of cancer remains a challenge. In order to extract useful gene information from cancer microarray data and reduce dimensionality, feature selection algorithms were systematically investigated in this study. Using a correlation-based feature selector combined with machine learning algorithms such as decision trees, naïve Bayes and support vector machines, we show that classification performance at least as good as published results can be obtained on acute leukemia and diffuse large B-cell lymphoma microarray data sets. We also demonstrate that a combined use of different classification and feature selection approaches makes it possible to select relevant genes with high confidence. This is also the first paper which discusses both computational and biological evidence for the involvement of zyxin in leukaemogenesis.

Algorithms↗

Bioinformatics challenges in proteomics.

A little after the genomic revolution had been celebrated, it seemed as if a competition began to found new -omics disciplines that ultimately all have the same goal, the understanding of biological function. There are many similar definitions for proteomics that can be summarized as follows: proteomics is a large-scale study of structure and function of proteins in an organism or cell. Importantly, the proteome is much more variable than the genome through its interactions with the genome and secondary modifications. It differs depending on the tissue and stage in life-cycle. Hence, proteomics is a very diverse discipline that uses a variety of experimental set-ups and targets in order to elucidate function. Its dissociation from other disciplines can only remain artificial. The bioinformatics applied to proteomics are equally varied. In this review we will focus mainly on a few areas of bioinformatics that seem to us as particularly noteworthy or characteristic for proteomics research, for example in 2DE analysis or mass spectrometry. Another important task of bioinformatics is the prediction of functional properties. We will summarize the approaches taken in order to predict protein networks, which are based on the extensive integration of several kinds of -omics data. We will give a short overview of a demanding field in computational biology, the analysis and prediction of protein 3D structures. In order to provide a broader perspective we will close this review with a generalized description of activities and databases in the realm of proteomics.

Animals↗

Iterative data analysis is the key for exhaustive analysis of peptide mass fingerprints from proteins separated by two-dimensional electrophoresis.

Peptide mass fingerprinting (PMF) is a powerful tool for identification of proteins separated by two-dimensional electrophoresis (2-DE). With the increase in sensitivity of peptide mass determination it becomes obvious that even spots looking well separated on a 2-DE gel may consist of several proteins. As a result the number of mass peaks in PMFs increased dramatically leaving many unassigned after a first database search. A number of these are caused by experiment-specific contaminants or by neighbor spots, as well as by additional proteins or post-translational modifications. To understand the complete protein composition of a spot we suggest an iterative procedure based on large numbers of PMFs, exemplified by PMFs of 480 Helicobacter pylori protein spots. Three key iterations were applied: (1) Elimination of contaminant mass peaks determined by MS-Screener (a software developed for this purpose) followed by reanalysis; (2) neighbor spot mass peak determination by cluster analysis, elimination from the peak list and repeated search; (3) re-evaluation of contaminant peaks. The quality of the identification was improved and spots previously unidentified were assigned to proteins. Eight additional spots were identified with this procedure, increasing the total number of identified spots to 455.

Cluster Analysis↗