Search PubMed⌕ Search

Biomedical subjects

Søren Brunak

Publications and source records attributed to Søren Brunak.

14 recordsLinked to original sources

Protein feature based identification of cell cycle regulated proteins in yeast.

DNA microarrays have been used extensively to identify cell cycle regulated genes in yeast; however, the overlap in the genes identified is surprisingly small. We show that certain protein features can be used to distinguish cell cycle regulated genes from other genes with high confidence (features include protein phosphorylation, glycosylation, subcellular location and instability/degradation). We demonstrate that co-expressed, periodic genes encode proteins which share combinations of features, and provide an overview of the proteome dynamics during the cycle. A large set of novel putative cell cycle regulated proteins were identified, many of which have no known function.

Cell Cycle↗

NESbase version 1.0: a database of nuclear export signals.

Protein export from the nucleus is often mediated by a Leucine-rich Nuclear Export Signal (NES). NESbase is a database of experimentally validated Leucine-rich NESs curated from literature. These signals are not annotated in databases such as SWISS-PROT, PIR or PROSITE. Each NESbase entry contains information of whether NES was shown to be necessary and/or sufficient for export, and whether the export was shown to be mediated by the export receptor CRM1. The compiled information was used to make a sequence logo of the Leucine-rich NESs, displaying the conservation of amino acids within a window of 25 residues. Surprisingly, only 36% of the sequences used for the logo fit the widely accepted NES consensus L-x(2,3)-[LIVFM]-x(2,3)-L-x-[LI]. The database is available online at http://www.cbs.dtu.dk/databases/NESbase/.

Active Transport, Cell Nucleus↗

Predicting proteasomal cleavage sites: a comparison of available methods.

The proteasome plays an essential role in the immune responses of vertebrates. By degrading intercellular proteins from self and non-self, the proteasome produces the majority of the peptides that are presented to cytotoxic T cells (CTL). There is accumulating evidence that the C-terminal, in particular, of CTL epitopes is cleaved precisely by the proteasome, whereas the N-terminal is produced with an extension, and later trimmed by peptidases in the cytoplasm and in the endoplasmic reticulum. Recently, three publicly available methods have been developed for prediction of the specificity of the proteasome. Here, we compare the performance of these methods on a large set of CTL epitopes. The best method, NetChop at www.cbs.dtu.dk/Services/NetChop, can capture approximately 70% of the C-termini correctly. This result suggests that the predictions can still be improved, particularly if more quantitative degradation data become available.

Algorithms↗

Reliable prediction of T-cell epitopes using neural networks with novel sequence representations.

In this paper we describe an improved neural network method to predict T-cell class I epitopes. A novel input representation has been developed consisting of a combination of sparse encoding, Blosum encoding, and input derived from hidden Markov models. We demonstrate that the combination of several neural networks derived using different sequence-encoding schemes has a performance superior to neural networks derived using a single sequence-encoding scheme. The new method is shown to have a performance that is substantially higher than that of other methods. By use of mutual information calculations we show that peptides that bind to the HLA A*0204 complex display signal of higher order sequence correlations. Neural networks are ideally suited to integrate such higher order correlations when predicting the binding affinity. It is this feature combined with the use of several neural networks derived from different and novel sequence-encoding schemes and the ability of the neural network to be trained on data consisting of continuous binding affinities that gives the new method an improved performance. The difference in predictive performance between the neural network methods and that of the matrix-driven methods is found to be most significant for peptides that bind strongly to the HLA molecule, confirming that the signal of higher order sequence correlation is most strongly present in high-binding peptides. Finally, we use the method to predict T-cell epitopes for the genome of hepatitis C virus and discuss possible applications of the prediction method to guide the process of rational vaccine design.

Amino Acid Sequence↗

A new non-linear normalization method for reducing variability in DNA microarray experiments.

BACKGROUND: Microarray data are subject to multiple sources of variation, of which biological sources are of interest whereas most others are only confounding. Recent work has identified systematic sources of variation that are intensity-dependent and non-linear in nature. Systematic sources of variation are not limited to the differing properties of the cyanine dyes Cy(5) and Cy(3) as observed in cDNA arrays, but are the general case for both oligonucleotide microarray (Affymetrix GeneChips) and cDNA microarray data. Current normalization techniques are most often linear and therefore not capable of fully correcting for these effects. RESULTS: We present here a simple and robust non-linear method for normalization using array signal distribution analysis and cubic splines. These methods compared favorably to normalization using robust local-linear regression (lowess). The application of these methods to oligonucleotide arrays reduced the relative error between replicates by 5-10% compared with a standard global normalization method. Application to cDNA arrays showed improvements over the standard method and over Cy(3)-Cy(5) normalization based on dye-swap replication. In addition, a set of known differentially regulated genes was ranked higher by the t-test. In either cDNA or Affymetrix technology, signal-dependent bias was more than ten times greater than the observed print-tip or spatial effects. CONCLUSIONS: Intensity-dependent normalization is important for both high-density oligonucleotide array and cDNA array data. Both the regression and spline-based methods described here performed better than existing linear methods when assessed on the variability of replicate arrays. Dye-swap normalization was less effective at Cy(3)-Cy(5) normalization than either regression or spline-based methods alone.

Cell Line↗

RNA secondary structure and squence conservation in C1 region of human immunodeficiency virus type 1 env gene.

We have analyzed amino acid, nucleotide sequence, and RNA secondary structure variability in the env gene of human immunodeficiency virus type (HIV-1). In applying algorithms for computing optimal RNA-folding patterns to a nonredundant data set of 178 env nucleotide sequences, we found a conserved RNA stem-loop structure in the first conserved (C1) region of the env gene. This detailed examination also revealed the known secondary structure conservation of the Rev-responsive element (RRE). This finding is also supported by a higher third position conservation of the translatable reading frame along these subregions. The typical folding of the C1 region consists of two isolated stem-loop structures. These highly conserved structures are likely to have a biological function. This assumption is supported by the conservation of the third position along the coding region of these structures. The third position retains a conservation level above what would be statistically expected.

Algorithms↗

Bias of purine stretches in sequenced chromosomes.

We examined more than 700 DNA sequences (full length chromosomes and plasmids) for stretches of purines (R) or pyrimidines (Y) and alternating YR stretches; such regions will likely adopt structures which are different from the canonical B-form. Since one turn of the DNA helix is roughly 10 bp, we measured the fraction of each genome which contains purine (or pyrimidine) tracts of lengths of 10 bp or longer (hereafter referred to as 'purine tracts'), as well as stretches of alternating pyrimidines/purine (pyr/pur tracts') of the same length. Using this criteria, a random sequence would be expected to contain 1.0% of purine tracts and also 1.0% of the alternating pyr/pur tracts. In the vast majority of cases, there are more purine tracts than would be expected from a random sequence, with an average of 3.5%, significantly larger than the expectation value. The fraction of the chromosomes containing pyr/pur tracts was slightly less than expected, with an average of 0.8%. One of the most surprising findings is a clear difference in the length distributions of the regions studied between prokaryotes and eukaryotes. Whereas short-range correlations can explain the length distributions in prokaryotes, in eukaryotes there is an abundance of long stretches of purines or alternating purine/pyrimidine tracts, which cannot be explained in this way; these sequences are likely to play an important role in eukaryotic chromosome organisation.

Animals↗

Continuum secondary structure captures protein flexibility.

The DSSP program assigns protein secondary structure to one of eight states. This discrete assignment cannot describe the continuum of thermal fluctuations. Hence, a continuous assignment is proposed. Technically, the continuum results from averaging over ten discrete DSSP assignments with different hydrogen bond thresholds. The final continuous assignment for a single NMR model successfully reflected the structural variations observed between all NMR models in the ensemble. The structural variations between NMR models were verified to correlate with thermal motion; these variations were captured by the continuous assignments. Because the continuous assignment reproduces the structural variation between many NMR models from one single model, functionally important variation can be extracted from a single X-ray structure. Thus, continuous assignments of secondary structure may affect future protein structure analysis, comparison, and prediction.

Computer Systems↗

Neural network predicts sequence of TP53 gene based on DNA chip.

UNLABELLED: We have trained an artificial neural network to predict the sequence of the human TP53 tumor suppressor gene based on a p53 GeneChip. The trained neural network uses as input the fluorescence intensities of DNA hybridized to oligonucleotides on the surface of the chip and makes between zero and four errors in the predicted 1300 bp sequence when tested on wild-type TP53 sequence. AVAILABILITY: The trained neural network is available for academic use by contacting steen@cbs.dtu.dk

Base Sequence↗

Prediction of proteasome cleavage motifs by neural networks.

We present a predictive method that can simulate an essential step in the antigen presentation in higher vertebrates, namely the step involving the proteasomal degradation of polypeptides into fragments which have the potential to bind to MHC Class I molecules. Proteasomal cleavage prediction algorithms published so far were trained on data from in vitro digestion experiments with constitutive proteasomes. As a result, they did not take into account the characteristics of the structurally modified proteasomes--often called immunoproteasomes--found in cells stimulated by gamma-interferon under physiological conditions. Our algorithm has been trained not only on in vitro data, but also on MHC Class I ligand data, which reflect a combination of immunoproteasome and constitutive proteasome specificity. This feature, together with the use of neural networks, a non-linear classification technique, make the prediction of MHC Class I ligand boundaries more accurate: 65% of the cleavage sites and 85% of the non-cleavage sites are correctly determined. Moreover, we show that the neural networks trained on the constitutive proteasome data learns a specificity that differs from that of the networks trained on MHC Class I ligands, i.e. the specificity of the immunoproteasome is different than the constitutive proteasome. The tools developed in this study in combination with a predictor of MHC and TAP binding capacity should give a more complete prediction of the generation and presentation of peptides on MHC Class I molecules. Here we demonstrate that such an approach produces an accurate prediction of the CTL the epitopes in HIV Nef. The method is available at www.cbs.dtu.dk/services/NetChop/.

Algorithms↗

Prediction of novel archaeal enzymes from sequence-derived features.

The completely sequenced archaeal genomes potentially encode, among their many functionally uncharacterized genes, novel enzymes of biotechnological interest. We have developed a prediction method for detection and classification of enzymes from sequence alone (available at http://www.cbs.dtu.dk/services/ArchaeaFun/). The method does not make use of sequence similarity; rather, it relies on predicted protein features like cotranslational and posttranslational modifications, secondary structure, and simple physical/chemical properties.

Amino Acid Motifs↗

Clustering patterns of cytotoxic T-lymphocyte epitopes in human immunodeficiency virus type 1 (HIV-1) proteins reveal imprints of immune evasion on HIV-1 global variation.

The human cytotoxic T-lymphocyte (CTL) response to human immunodeficiency virus type 1 (HIV-1) has been intensely studied, and hundreds of CTL epitopes have been experimentally defined, published, and compiled in the HIV Molecular Immunology Database. Maps of CTL epitopes on HIV-1 protein sequences reveal that defined epitopes tend to cluster. Here we integrate the global sequence and immunology databases to systematically explore the relationship between HIV-1 amino acid sequences and CTL epitope distributions. CTL responses to five HIV-1 proteins, Gag p17, Gag p24, reverse transcriptase (RT), Env, and Nef, have been particularly well characterized in the literature to date. Through comparing CTL epitope distributions in these five proteins to global protein sequence alignments, we identified distinct characteristics of HIV amino acid sequences that correlate with CTL epitope localization. First, experimentally defined HIV CTL epitopes are concentrated in relatively conserved regions. Second, the highly variable regions that lack epitopes bear cumulative evidence of past immune escape that may make them relatively refractive to CTLs: a paucity of predicted proteasome processing sites and an enrichment for amino acids that do not serve as C-terminal anchor residues. Finally, CTL epitopes are more highly concentrated in alpha-helical regions of proteins. Based on amino acid sequence characteristics, in a blinded fashion, we predicted regions in HIV regulatory and accessory proteins that would be likely to contain CTL epitopes; these predictions were then validated by comparison to new sets of experimentally defined epitopes in HIV-1 Rev, Tat, Vif, and Vpr.

Amino Acid Sequence↗

Sigma A recognition sites in the Bacillus subtilis genome.

A hidden Markov model of sigma(A) RNA polymerase cofactor recognition sites in Bacillus subtilis, containing either the common or the extended -10 motifs, has been constructed based on experimentally verified sigma(A) recognition sites. This work suggests that more information exists at the initiation site of transcription in both types of promoters than previously thought. When tested on the entire B. subtilis genome, the model predicts that approximately half of the sigma(A) recognition sites are of the extended type. Some of the response-regulator aspartate phosphatases were among the predictions of promoters containing extended sites. The expression of rapA and rapB was confirmed by site-directed mutagenesis to depend on the extended -10 region.

Bacillus subtilis↗