Search PubMed⌕ Search

Biomedical subjects

S Brunak

Publications and source records attributed to S Brunak.

At least 55 records · Page 3Linked to original sources

NetOglyc: prediction of mucin type O-glycosylation sites based on sequence context and surface accessibility.

The specificities of the UDP-GalNAc:polypeptide Nacetylgalactosaminyltransferases which link the carbohydrate GalNAc to the side-chain of certain serine and threonine residues in mucin type glycoproteins, are presently unknown. The specificity seems to be modulated by sequence context, secondary structure and surface accessibility. The sequence context of glycosylated threonines was found to differ from that of serine, and the sites were found to cluster. Non-clustered sites had a sequence context different from that of clustered sites. Charged residues were disfavoured at position -1 and +3. A jury of artificial neural networks was trained to recognize the sequence context and surface accessibility of 299 known and verified mucin type O-glycosylation sites extracted from O-GLYCBASE. The cross-validated NetOglyc network system correctly found 83% of the glycosylated and 90% of the non-glycosylated serine and threonine residues in independent test sets, thus proving more accurate than matrix statistics and vector projection methods. Predictions of O-glycosylation sites in the envelope glycoprotein gp120 from the primate lentiviruses HIV-1, HIV-2 and SIV are presented. The most conserved O-glycosylation signals in these evolutionary-related glycoproteins were found in their first hypervariable loop, V1. However, the strain variation for HIV-1 gp120 was significant. A computer server, available through WWW or E-mail, has been developed for prediction of mucin type O-glycosylation sites in proteins based on the amino acid sequence. The server addresses are http://www.cbs.dtu.dk/services/NetOGlyc/ and netOglyc@cbs.dtu.dk.

Algorithms↗

Computational applications of DNA structural scales.

We study from a computational standpoint several different physical scales associated with structural features of DNA sequences, including dinucleotide scales such as base stacking energy and propeller twist, and trinucleotide scales such as bendability and nucleosome positioning. We show that these scales provide an alternative or complementary compact representation of DNA sequences. As an example we construct a strand invariant representation of DNA sequences. The scales can also be used to analyze and discover new DNA structural patterns, especially in combinations with hidden Markov models (HMMs). The scales are applied to HMMs of human promoter sequences revealing a number of significant differences between regions upstream and downstream of the transcriptional start point. Finally we show, with some qualifications, that such scales are by and large independent, and therefore complement each other.

Artificial Intelligence↗

A branch point consensus from Arabidopsis found by non-circular analysis allows for better prediction of acceptor sites.

Little knowledge exists about branch points in plants; it has even been claimed that plant introns lack conserved branch point sequences similar to those found in vertebrate introns. A putative branch point consensus sequence for Arabidopsis thaliana resembling the well known metazoan consensus sequence has been proposed, but this is based on search of sequences similar to those in yeast and metazoa. Here we present a novel consensus sequence found by a non-circular approach. A hidden Markov model with a fixed A nucleotide was trained on sequences upstream of the acceptor site. The consensus found by the Markov model shares features with the metazoan consensus, but differs in its details from the consensus proposed earlier. Despite the fact that branch point consensus sequences in plants are weak, we show that a prediction scheme incorporating them leads to a substantial improvement in the recognition of true acceptor sites; the false positive rate being reduced by a factor of 2. We take this as an indication that the consensus found here is the genuine one and that the branch point does play a role in the proper recognition of the acceptor site in plants.

Arabidopsis↗

Protein folding and wring resonances.

The polypeptide chain of a protein is shown to obey topological constraints which enable long range excitations in the form of wring modes of the protein backbone. Wring modes of proteins of specific lengths can therefore resonate with molecular modes present in the cell. It is suggested that protein folding takes place when the amplitude of a wring excitation becomes so large that it is energetically favorable to bend the protein backbone. The condition under which such structural transformations can occur is found, and it is shown that both cold and hot denaturation (the unfolding of proteins) are natural consequences of the suggested wring mode model. Native (folded) proteins are found to possess an intrinsic standing wring mode.

Journal Article↗

O-GLYCBASE version 2.0: a revised database of O-glycosylated proteins.

O-GLYCBASE is an updated database of information on glycoproteins and their O-linked glycosylation sites. Entries are compiled and revised from the literature, and from the SWISS-PROT database. Entries include information about species, sequence, glycosylation sites and glycan type. O-GLYCBASE is now fully cross-referenced to the SWISS-PROT, PIR, PROSITE, PDB, EMBL, HSSP, LISTA and MIM databases. Compared with version 1.0 the number of entries have increased by 34%. Revision of the O-glycan assignment was performed on 20% of the entries. Sequence logos displaying the acceptor specificity patterns for the GalNAc, mannose and GlcNAc transferases are shown. The O-GLYCBASE database is available through WWW or by anonymous FTP.

Algorithms↗

Molecular wring resonances in chain molecules.

It is shown that the eigenfrequency of collective twist excitations in chain molecules can be in the megahertz and gigahertz range. Accordingly, resonance states can be obtained at specific frequencies, and phenomena that involve structural properties can take place. Chain molecules can alter their conformation and their ability to function, and a breaking of the chain can result. It is suggested that this phenomenon forms the basis for effects caused by the interaction of microwaves and biomolecules, e.g., microwave assisted hydrolysis of chain molecules.

Biopolymers↗

Prediction of N-terminal protein sorting signals.

Recently, neural networks have been applied to a widening range of problems in molecular biology. An area particularly suited to neural-network methods is the identification of protein sorting signals and the prediction of their cleavage sites, as these functional units are encoded by local, linear sequences of amino acids rather than global 3D structures.

Chloroplasts↗

Coherent topological phenomena in protein folding.

A theory is presented for coherent topological phenomena in protein dynamics with implications for protein folding and stability. We discuss the relationship to the writhing number used in knot diagrams of DNA. The winding state defines a long-range order along the backbone of a protein with long-range excitations, 'wring' modes, that play an important role in protein denaturation and stability. Energy can be pumped into these excitations, either thermally or by an external force.

Mathematics↗

Displaying the information contents of structural RNA alignments: the structure logos.

MOTIVATION: We extend the standard 'Sequence Logo' method of Schneider and Stevens (Nucleic Acids Res., 18, 6097-6100, 1990) to incorporate prior frequencies on the bases, allow for gaps in the alignments, and indicate the mutual information of base-paired regions in RNA. RESULTS: Given an alignment of RNA sequences with the base pairings indicated, the program will calculate the information at each position, including the mutual information of the base pairs, and display the results in a 'Structure Logo'. Alignments without base pairing can also be displayed in a 'Sequence Logo', but still allowing gaps and incorporating prior frequencies if desired. AVAILABILITY: The code is available from, and an Internet server can be used to run the program at, http://www.cbs.dtu.dk/gorodkin/appl/slogo. html.

Algorithms↗

Identification of prokaryotic and eukaryotic signal peptides and prediction of their cleavage sites.

We have developed a new method for the identification of signal peptides and their cleavage sites based on neural networks trained on separate sets of prokaryotic and eukaryotic sequence. The method performs significantly better than previous prediction schemes and can easily be applied on genome-wide data sets. Discrimination between cleaved signal peptides and uncleaved N-terminal signal-anchor sequences is also possible, though with lower precision. Predictions can be made on a publicly available WWW server.

Algorithms↗

Protein distance constraints predicted by neural networks and probability density functions.

We predict interatomic Calpha distances by two independent data driven methods. The first method uses statistically derived probability distributions of the pairwise distance between two amino acids, whilst the latter method consists of a neural network prediction approach equipped with windows taking the context of the two residues into account. These two methods are used to predict whether distances in independent test sets were above or below given thresholds. We investigate which distance thresholds produce the most information-rich constraints and, in turn, the optimal performance of the two methods. The predictions are based on a data set derived using a new threshold which defines when sequence similarity implies structural similarity. We show that distances in proteins are predicted more accurately by neural networks than by probability density functions. We show that the accuracy of the predictions can be further increased by using sequence profiles. A threading method based on the predicted distances is presented. A homepage with software, predictions and data related to this paper is available at http://www.cbs.dtu.dk/services/CPHmodels/.

Amino Acids↗

Naturally occurring nucleosome positioning signals in human exons and introns.

We describe the structural implications of a periodic pattern found in human exons and introns by hidden Markov models. We show that exons (besides the reading frame) have a specific sequential structure in the form of a pattern with triplet consensus non-T(A/T)G, and a minimal periodicity of roughly ten nucleotides. The periodic pattern is also present in intron sequences, although the strength per nucleotide is weaker. Using two independent profile methods based on triplet bendability parameters from DNase I experiments and nucleosome positioning data, we show that the pattern in multiple alignments of internal exon and intron sequences corresponds to a periodic "in phase" bending potential towards the major groove of the DNA. The nucleosome positioning data show that the consensus triplets (and their complements) have a preference for locations on a bent double helix where the major groove faces inward and is compressed. The in-phase triplets are located adjacent to GCC/GGC triplets known to have the strongest bias in their positioning on the nuclesome. Analysis of mRNA sequences encoding proteins with known tertiary structure exclude the possibility that the pattern is a consequence of the previously well-known periodicity caused by the encoding of alpha-helices in proteins. Finally, we discuss the relation between the bending potential of coding and non-coding regions and its impact on the translational positioning of nucleosomes and the recognition of genes by the transcriptional machinery.

Base Sequence↗

Splice site prediction in Arabidopsis thaliana pre-mRNA by combining local and global sequence information.

Artificial neural networks have been combined with a rule based system to predict intron splice sites in the dicot plant Arabidopsis thaliana. A two step prediction scheme, where a global prediction of the coding potential regulates a cutoff level for a local prediction of splice sites, is refined by rules based on splice site confidence values, prediction scores, coding context and distances between potential splice sites. In this approach, the prediction of splice sites mutually affect each other in a non-local manner. The combined approach drastically reduces the large amount of false positive splice sites normally haunting splice site prediction. An analysis of the errors made by the networks in the first step of the method revealed a previously unknown feature, a frequent T-tract prolongation containing cryptic acceptor sites in the 5' end of exons. The method presented here has been compared with three other approaches, GeneFinder, Gene-Mark and Grail. Overall the method presented here is an order of magnitude better. We show that the new method is able to find a donor site in the coding sequence for the jelly fish Green Fluorescent Protein, exactly at the position that was experimentally observed in A.thaliana transformants. Predictions for alternatively spliced genes are also presented, together with examples of genes from other dicots, monocots and algae. The method has been made available through electronic mail (NetPlantGene@cbs.dtu.dk), or the WWW at http://www.cbs.dtu.dk/NetPlantGene.html

Algorithms↗

Cleaning the GenBank Arabidopsis thaliana data set.

Data driven computational biology relies on the large quantities of genomic data stored in international sequence data banks. However, the possibilities are drastically impaired if the stored data is unreliable. During a project aiming to predict splice sites in the dicot Arabidopsis thaliana, we extracted a data set from the A.thaliana entries in GenBank. A number of simple 'sanity' checks, based on the nature of the data, revealed an alarmingly high error rate. More than 15% of the most important entries extracted did contain erroneous information. In addition, a number of entries had directly conflicting assignments of exons and introns, not stemming from alternative splicing. In a few cases the errors are due to mere typographical misprints, which may be corrected by comparison to the original papers, but errors caused by wrong assignments of splice sites from experimental data are the most common. It is proposed that the level of error correction should be increased and that gene structure sanity checks should be incorporated--also at the submitter level--to avoid or reduce the problem in the future. A non-redundant and error corrected subset of the data for A.thaliana is made available through anonymous FTP.

Algorithms↗

O-GLYCBASE: a revised database of O-glycosylated proteins.

O-GLYCBASE is a comprehensive database of information on glycoproteins and their O-linked glycosylation sites. Entries are compiled and revised from the SWISS-PROT and PIR databases as well as directly from recently published reports. Nineteen percent of the entries extracted from the databases needed revision with respect to O-linked glycosylation. Entries include information about species, sequence, glycosylation site and glycan type, and are fully referenced. Sequence logos displaying the acceptor specificity for the GaINAc transferase are shown. A neural network method for prediction of mucin type O-glycosylation sites in mammalian glycoproteins exclusively from the primary sequence is made available by E-mail or WWW. The O-GLYCBASE database is also available electronically through our WWW server or by anonymous FTP.

Amino Acid Sequence↗

Defining a similarity threshold for a functional protein sequence pattern: the signal peptide cleavage site.

When preparing data sets of amino acid or nucleotide sequences it is necessary to exclude redundant or homologous sequences in order to avoid overestimating the predictive performance of an algorithm. For some time methods for doing this have been available in the area of protein structure prediction. We have developed a similar procedure based on pair-wise alignments for sequences with functional sites. We show how a correlation coefficient between sequence similarity and functional homology can be used to compare the efficiency of different similarity measures and choose a nonarbitrary threshold value for excluding redundant sequences. The impact of the choice of scoring matrix used in the alignments is examined. We demonstrate that the parameter determining the quality of the correlation is the relative entropy of the matrix, rather than the assumed (PAM or identity) substitution mode. Results are presented for the case of prediction of cleavage sites in signal peptides. By inspection of the false positives, several errors in the database were found. The procedure presented may be used as a general outline for finding a problem-specific similarity measure and threshold value for analysis of other functional amino acid or nucleotide sequence patterns.

Algorithms↗

Prediction of the secondary structure of HIV-1 gp120.

The secondary structure of HIV-1 gp120 was predicted using multiple alignment and a combination of two independent methods based on neural network and nearest-neighbor algorithms. The methods agreed on the secondary structure for 80% of the residues in BH10 gp120. Six helices were predicted in HIV strain BH10 gp120, as well as in 27 other HIV-1 strains examined. Two helical segments were predicted in regions displaying profound sequence variation, one in a region suggested to be critical for CD4 binding. The predicted content of helix, beta-strand, and coil was consistent with estimates from Fourier transform infrared spectroscopy. The predicted secondary structure of gp120 compared well with data from NMR analysis of synthetic peptides from the V3 loop and the C4 region. As a first step towards modeling the tertiary structure of gp120, the predicted secondary structure may guide the design of future HIV subunit vaccine candidates.

Algorithms↗