Search PubMed⌕ Search

Biomedical subjects

S Brunak

Publications and source records attributed to S Brunak.

At least 19 recordsLinked to original sources

Genome organisation and chromatin structure in Escherichia coli.

We have analysed the complete sequence of the Escherichia coli K12 isolate MG1655 genome for chromatin-associated protein binding sites, and compared the predicted location of predicted sites with experimental expression data from 'DNA chip' experiments. Of the dozen proteins associated with chromatin in E. coli, only three have been shown to have significant binding preferences: integration host factor (IHF) has the strongest binding site preference, and FIS sites show a weak consensus, and there is no clear consensus site for binding of the H-NS protein. Using hidden Markov models (HMMs), we predict the location of 608 IHF sites, scattered throughout the genome. A subset of the IHF sites associated with repeats tends to be clustered around the origin of replication. We estimate there could be roughly 6000 FIS sites in E. coli, and the sites tend to be localised in two regions flanking the replication termini. We also show that the regions upstream of genes regulated by H-NS are more curved and have a higher AT content than regions upstream of other genes. These regions in general would also be localised near the replication terminus.

Bacterial Proteins↗

Identification and design of p53-derived HLA-A2-binding peptides with increased CTL immunogenicity.

The replacement of a suboptimal amino acid in a primary anchor position with an optimal residue improves human leucocyte antigen (HLA) binding and immunogenicity, while maintaining cytotoxic T lymphocyte (CTL) specificity. Using a neural network capable of performing quantitative predictions of peptide binding to HLA-A2 molecules, we identified three p53 protein-derived nonamer peptides with intermediate binding owing to suboptimal amino acids in the P2 anchor position. These peptides were synthesized along with the corresponding analogs, where the natural P2 residue had been replaced with the optimal leucine residue. All three modified peptides bound to and more efficiently stabilized HLA-A2 molecules than the corresponding nonmodified peptides. The HLA-A2 transgenic mice were used for immunization. Two of the epitopes were more immunogenic in their modified than in their natural versions. The CTLs raised against the modified peptides efficiently killed the target cells pulsed with the corresponding native peptide. In terms of sensitizing the targets cells for the CTL killing, the modified peptides were more efficient than native peptides. Finally, the CTLs induced by modified peptide killed HLA-A2 transgenic mouse fibrosarcoma cells transfected with human p53 DNA. The data suggest that modified self-peptides derived from overexpressed tumour-associated proteins can be used in vaccine development against cancer, and that quantitative predictions of HLA binding is of value in the rational selection and improvement of target epitopes recognized by CTLs.

Amino Acid Sequence↗

Flexibility of the genetic code with respect to DNA structure.

MOTIVATION: The primary function of DNA is to carry genetic information through the genetic code. DNA, however, contains a variety of other signals related, for instance, to reading frame, codon bias, pairwise codon bias, splice sites and transcription regulation, nucleosome positioning and DNA structure. Here we study the relationship between the genetic code and DNA structure and address two questions. First, to which degree does the degeneracy of the genetic code and the acceptable amino acid substitution patterns allow for the superimposition of DNA structural signals to protein coding sequences? Second, is the origin or evolution of the genetic code likely to have been constrained by DNA structure? RESULTS: We develop an index for code flexibility with respect to DNA structure. Using five different di- or tri-nucleotide models of sequence-dependent DNA structure, we show that the standard genetic code provides a fair level of flexibility at the level of broad amino acid categories. Thus the code generally allows for the superimposition of any structural signal on any protein-coding sequence, through amino acid substitution. The flexibility observed at the level of single amino acids allows only for the superimposition of punctual and loosely positioned signals to conserved amino acid sequences. The degree of flexibility of the genetic code is low or average with respect to several classes of alternative codes. This result is consistent with the view that DNA structure is not likely to have played a significant role in the origin and evolution of the genetic code.

Amino Acids↗

Prediction of protein secondary structure at 80% accuracy.

Secondary structure prediction involving up to 800 neural network predictions has been developed, by use of novel methods such as output expansion and a unique balloting procedure. An overall performance of 77.2%-80.2% (77.9%-80.6% mean per-chain) for three-state (helix, strand, coil) prediction was obtained when evaluated on a commonly used set of 126 protein chains. The method uses profiles made by position-specific scoring matrices as input, while at the output level it predicts on three consecutive residues simultaneously. The predictions arise from tenfold, cross validated training and testing of 1032 protein sequences, using a scheme with primary structure neural networks followed by structure filtering neural networks. With respect to blind prediction, this work is preliminary and awaits evaluation by CASP4.

Neural Networks, Computer↗

Predicting subcellular localization of proteins based on their N-terminal amino acid sequence.

A neural network-based tool, TargetP, for large-scale subcellular location prediction of newly identified proteins has been developed. Using N-terminal sequence information only, it discriminates between proteins destined for the mitochondrion, the chloroplast, the secretory pathway, and "other" localizations with a success rate of 85% (plant) or 90% (non-plant) on redundancy-reduced test sets. From a TargetP analysis of the recently sequenced Arabidopsis thaliana chromosomes 2 and 4 and the Ensembl Homo sapiens protein set, we estimate that 10% of all plant proteins are mitochondrial and 14% chloroplastic, and that the abundance of secretory proteins, in both Arabidopsis and Homo, is around 10%. TargetP also predicts cleavage sites with levels of correctly predicted sites ranging from approximately 40% to 50% (chloroplastic and mitochondrial presequences) to above 70% (secretory signal peptides). TargetP is available as a web-server at http://www.cbs.dtu.dk/services/TargetP/.

Amino Acid Sequence↗

A DNA structural atlas for Escherichia coli.

We have performed a computational analysis of DNA structural features in 18 fully sequenced prokaryotic genomes using models for DNA curvature, DNA flexibility, and DNA stability. The structural values that are computed for the Escherichia coli chromosome are significantly different from (and generally more extreme than) that expected from the nucleotide composition. To aid this analysis, we have constructed tools that plot structural measures for all positions in a long DNA sequence (e.g. an entire chromosome) in the form of color-coded wheels (http://www.cbs.dtu. dk/services/GenomeAtlas/). We find that these "structural atlases" are useful for the discovery of interesting features that may then be investigated in more depth using statistical methods. From investigation of the E. coli structural atlas, we discovered a genome-wide trend, where an extended region encompassing the terminus displays a high of level curvature, a low level of flexibility, and a low degree of helix stability. The same situation is found in the distantly related Gram-positive bacterium Bacillus subtilis, suggesting that the phenomenon is biologically relevant. Based on a search for long DNA segments where all the independent structural measures agree, we have found a set of 20 regions with identical and very extreme structural properties. Due to their strong inherent curvature, we suggest that these may function as topological domain boundaries by efficiently organizing plectonemically supercoiled DNA. Interestingly, we find that in practically all the investigated eubacterial and archaeal genomes, there is a trend for promoter DNA being more curved, less flexible, and less stable than DNA in coding regions and in intergenic DNA without promoters. This trend is present regardless of the absolute levels of the structural parameters, and we suggest that this may be related to the requirement for helix unwinding during initiation of transcription, or perhaps to the previously observed location of promoters at the apex of plectonemically supercoiled DNA. We have also analyzed the structural similarities between groups of genes by clustering all RNA and protein-encoding genes in E. coli, based on the average structural parameters. We find that most ribosomal genes (protein-encoding as well as rRNA genes) cluster together, and we suggest that DNA structure may play a role in the transcription of these highly expressed genes.

Bacterial Proteins↗

Structural analysis of DNA sequence: evidence for lateral gene transfer in Thermotoga maritima.

The recently published complete DNA sequence of the bacterium Thermotoga maritima provides evidence, based on protein sequence conservation, for lateral gene transfer between Archaea and Bacteria. We introduce a new method of periodicity analysis of DNA sequences, based on structural parameters, which brings independent evidence for the lateral gene transfer in the genome of T.maritima. The structural analysis relates the Archaea-like DNA sequences to the genome of Pyrococcus horikoshii. Analysis of 24 complete genomic DNA sequences shows different periodicity patterns for organisms of different origin. The typical genomic periodicity for Bacteria is 11 bp whilst it is 10 bp for Archaea. Eukaryotes have more complex spectra but the dominant period in the yeast Saccharomyces cerevisiae is 10.2 bp. These periodicities are most likely reflective of differences in chromatin structure.

Chromatin↗

Assessing the accuracy of prediction algorithms for classification: an overview.

We provide a unified overview of methods that currently are widely used to assess the accuracy of prediction algorithms, from raw percentages, quadratic error measures and other distances, and correlation coefficients, and to information theoretic measures such as relative entropy and mutual information. We briefly discuss the advantages and disadvantages of each approach. For classification tasks, we derive new learning algorithms for the design of prediction systems by directly optimising the correlation coefficient. We observe and prove several results relating sensitivity and specificity of optimal systems. While the principles are general, we illustrate the applicability on specific problems such as protein secondary structure and signal peptide prediction.

Algorithms↗

env sequences of simian immunodeficiency viruses from chimpanzees in Cameroon are strongly related to those of human immunodeficiency virus group N from the same geographic area.

Human immunodeficiency virus type 1 (HIV-1) group N from Cameroon is phylogenetically close, in env, to the simian immunodeficiency virus (SIV) cpz-gab from Gabon and SIVcpz-US of unknown geographic origin. We screened 29 wild-born Cameroonian chimpanzees and found that three (Cam3, Cam4, and Cam5) were positive for HIV-1 by Western blotting. Mitochondrial DNA sequence analysis demonstrated that Cam3 and Cam5 belonged to Pan troglodytes troglodytes and that Cam4 belonged to P. t. vellerosus. Genetic analyses of the viruses together with serological data demonstrated that at least one of the two P. t. troglodytes chimpanzees (Cam5) was infected in the wild, and revealed a horizontal transmission between Cam3 and Cam4. These data confirm that P. t. troglodytes is a natural host for HIV-1-related viruses. Furthermore, they show that SIVcpz can be transmitted in captivity, from one chimpanzee subspecies to another. All three SIVcpz-cam viruses clustered with HIV-1 N in env. The full Cam3 SIVcpz genome sequence showed a very close phylogenetic relationship with SIVcpz-US, a virus identified in a P. t. troglodytes chimpanzee captured nearly 40 years earlier. Like SIVcpz-US, SIVcpz-cam3 was closely related to HIV-1 N in env, but not in pol, supporting the hypothesis that HIV-1 N results from a recombination event. SIVcpz from chimpanzees born in the wild in Cameroon are thus strongly related in env to HIV-1 N from Cameroon, demonstrating the geographic coincidence of these human and simian viruses and providing a further strong argument in favor of the origin of HIV-1 being in chimpanzees.

Animals↗

Matching protein beta-sheet partners by feedforward and recurrent neural networks.

Predicting the secondary structure (alpha-helices, beta-sheets, coils) of proteins is an important step towards understanding their three dimensional conformations. Unlike alpha-helices that are built up from one contiguous region of the polypeptide chain, beta-sheets are more complex resulting from a combination of two or more disjoint regions. The exact nature of these long distance interactions remains unclear. Here we introduce two neural-network based methods for the prediction of amino acid partners in parallel as well as anti-parallel beta-sheets. The neural architectures predict whether two residues located at the center of two distant windows are paired or not in a beta-sheet structure. Variations on these architecture, including also profiles and ensembles, are trained and tested via five-fold cross validation using a large corpus of curated data. Prediction on both coupled and non-coupled residues currently approaches 84% accuracy, better than any previously reported method.

Animals↗

Sequence and structure-based prediction of eukaryotic protein phosphorylation sites.

Protein phosphorylation at serine, threonine or tyrosine residues affects a multitude of cellular signaling processes. How is specificity in substrate recognition and phosphorylation by protein kinases achieved? Here, we present an artificial neural network method that predicts phosphorylation sites in independent sequences with a sensitivity in the range from 69 % to 96 %. As an example, we predict novel phosphorylation sites in the p300/CBP protein that may regulate interaction with transcription factors and histone acetyltransferase activity. In addition, serine and threonine residues in p300/CBP that can be modified by O-linked glycosylation with N-acetylglucosamine are identified. Glycosylation may prevent phosphorylation at these sites, a mechanism named yin-yang regulation. The prediction server is available on the Internet at http://www.cbs.dtu.dk/services/NetPhos/or via e-mail to NetPhos@cbs. dtu.dk.

Amino Acid Motifs↗

The biology of eukaryotic promoter prediction--a review.

Computational prediction of eukaryotic promoters from the nucleotide sequence is one of the most attractive problems in sequence analysis today, but it is also a very difficult one. Thus, current methods predict in the order of one promoter per kilobase in human DNA, while the average distance between functional promoters has been estimated to be in the range of 30-40 kilobases. Although it is conceivable that some of these predicted promoters correspond to cryptic initiation sites that are used in vivo, it is likely that most are false positives. This suggests that it is important to carefully reconsider the biological data that forms the basis of current algorithms, and we here present a review of data that may be useful in this regard. The review covers the following topics: (1) basal transcription and core promoters, (2) activated transcription and transcription factor binding sites, (3) CpG islands and DNA methylation, (4) chromosomal structure and nucleosome modification, and (5) chromosomal domains and domain boundaries. We discuss the possible lessons that may be learned, especially with respect to the wealth of information about epigenetic regulation of transcription that has been appearing in recent years.

Chromosomes↗

PhosphoBase, a database of phosphorylation sites: release 2.0.

PhosphoBase contains information about phosphorylated residues in proteins and data about peptide phosphorylation by a variety of protein kinases. The data are collected from literature and compiled into a common format. The current release of PhosphoBase (October 1998, version 2.0) comprises 414 phosphoprotein entries covering 1052 phosphorylatable serine, threonine and tyrosine residues. The kinetic data from peptide phosphorylation assays for approximately 330 oligopeptides is also included. The database entries are cross-referenced to the corresponding records in the Swiss-Prot protein database and literature references are linked to MedLine records. PhosphoBase is available via the WWW at http://www.cbs.dtu. dk/databases/PhosphoBase/

Animals↗

O-GLYCBASE version 4.0: a revised database of O-glycosylated proteins.

O-GLYCBASE is a database of glycoproteins with O-linked glycosylation sites. Entries with at least one experimentally verified O-glycosylation site have been compiled from protein sequence databases and literature. Each entry contains information about the glycan involved, the species, sequence, a literature reference and http-linked cross-references to other databases. Version 4.0 contains 179 protein entries, an approximate 15% increase over the last version. Sequence logos representing the acceptor specificity patterns for GalNAc, GlcNAc, mannosyl and xylosyl transferases are shown. The O-GLYCBASE database is available through the WWW at http://www.cbs.dtu.dk/databases/OGLYCBASE/

Acetylgalactosamine↗

Structural basis for triplet repeat disorders: a computational analysis.

MOTIVATION: Over a dozen major degenerative disorders, including myotonic distrophy, Huntington's disease and fragile X syndrome, result from unstable expansions of particular trinucleotides. Remarkably, only some of all the possible triplets, namely CAG/CTG, CGG/CCG and GAA/TTC, have been associated with the known pathological expansions. This raises some basic questions at the DNA level. Why do particular triplets seem to be singled out? What is the mechanism for their expansion and how does it depend on the triplet itself? Could other triplets or longer repeats be involved in other diseases? RESULTS: Using several different computational models of DNA structure, we show that the triplets involved in the pathological repeats generally fall into extreme classes. Thus, CAG/CTG repeats are particularly flexible, whereas GCC, CGG and GAA repeats appear to display both flexible and rigid (but curved) characteristics depending on the method of analysis. The fact that (1) trinucleotide repeats often become increasingly unstable when they exceed a length of approximately 50 repeats, and (2) repeated 12-mers display a similar increase in instability above 13 repeats, together suggest that approximately 150 bp is a general threshold length for repeat instability. Since this is about the length of DNA wrapped up in a single nucleosome core particle, we speculate that chromatin structure may play an important role in the expansion mechanism. We furthermore suggest that expansion of a dodecamer repeat, which we predict to have very high flexibility, may play a role in the pathogenesis of the neurodegenerative disorder multiple system atrophy (MSA). CONTACT: pfbaldi@ics.uci.edu, yves@netid.com, brunak@cbs.dtu.dk, gorm@cbs.dtu.dk.

Anticipation, Genetic↗

Exploiting the past and the future in protein secondary structure prediction.

MOTIVATION: Predicting the secondary structure of a protein (alpha-helix, beta-sheet, coil) is an important step towards elucidating its three-dimensional structure, as well as its function. Presently, the best predictors are based on machine learning approaches, in particular neural network architectures with a fixed, and relatively short, input window of amino acids, centered at the prediction site. Although a fixed small window avoids overfitting problems, it does not permit capturing variable long-rang information. RESULTS: We introduce a family of novel architectures which can learn to make predictions based on variable ranges of dependencies. These architectures extend recurrent neural networks, introducing non-causal bidirectional dynamics to capture both upstream and downstream information. The prediction algorithm is completed by the use of mixtures of estimators that leverage evolutionary information, expressed in terms of multiple alignments, both at the input and output levels. While our system currently achieves an overall performance close to 76% correct prediction--at least comparable to the best existing systems--the main emphasis here is on the development of new algorithmic ideas. AVAILABILITY: The executable program for predicting protein secondary structure is available from the authors free of charge. CONTACT: pfbaldi@ics.uci.edu, gpollast@ics.uci.edu, brunak@cbs.dtu.dk, paolo@dsi.unifi.it.

Algorithms↗

MatrixPlot: visualizing sequence constraints.

UNLABELLED: MatrixPlot is a program for making high-quality matrix plots, such as mutual information plots of sequence alignments and distance matrices of sequences with known three-dimensional coordinates. The user can add information about the sequences (e.g. a sequence logo profile) along the edges of the plot, as well as zoom in on any region in the plot. AVAILABILITY: MatrixPlot can be obtained on request, and can also be accessed online at http://www. cbs.dtu.dk/services/MatrixPlot. CONTACT: gorodkin@cbs.dtu.dk

Nucleic Acids↗

Scanning the available Dictyostelium discoideum proteome for O-linked GlcNAc glycosylation sites using neural networks.

Dictyostelium discoideum has been suggested as a eukaryotic model organism for glycobiology studies. Presently, the characteristics of acceptor sites for the N-acetylglucosaminyl-transferases in Dictyostelium discoideum, which link GlcNAc in an alpha linkage to hydroxyl residues, are largely unknown. This motivates the development of a species specific method for prediction of O-linked GlcNAc glycosylation sites in secreted and membrane proteins of D. discoideum. The method presented here employs a jury of artificial neural networks. These networks were trained to recognize the sequence context and protein surface accessibility in 39 experimentally determined O-alpha-GlcNAc sites found in D. discoideum glycoproteins expressed in vivo. Cross-validation of the data revealed a correlation in which 97% of the glycosylated and nonglycosylated sites were correctly identified. Based on the currently limited data set, an abundant periodicity of two (positions-3, -1, +1, +3, etc.) in Proline residues alternating with hydroxyl amino acids was observed upstream and downstream of the acceptor site. This was a consequence of the spacing of the glycosylated residues themselves which were peculiarly found to be situated only at even positions with respect to each other, indicating that these may be located within beta-strands. The method has been used for a rapid and ranked scan of the fraction of the Dictyostelium proteome available in public databases, remarkably 25-30% of which were predicted glycosylated. The scan revealed acceptor sites in several proteins known experimentally to be O-glycosylated at unmapped sites. The available proteome was classified into functional and cellular compartments to study any preferential patterns of glycosylation. A sequence based prediction server for GlcNAc O-glycosylations in D. discoideum proteins has been made available through the WWW at http://www.cbs.dtu.dk/services/DictyOGlyc/ and via E-mail to DictyOGlyc@cbs.dtu.dk.

Algorithms↗