Search PubMed⌕ Search

Biomedical subjects

S Brunak

Publications and source records attributed to S Brunak.

At least 73 records · Page 4Linked to original sources

Protein structure and the sequential structure of mRNA: alpha-helix and beta-sheet signals at the nucleotide level.

A direct comparison of experimentally determined protein structures and their corresponding protein coding mRNA sequences has been performed. We examine whether real world data support the hypothesis that clusters of rare codons correlate with the location of structural units in the resulting protein. The degeneracy of the genetic code allows for a biased selection of codons which may control the translational rate of the ribosome, and may thus in vivo have a catalyzing effect on the folding of the polypeptide chain. A complete search for GenBank nucleotide sequences coding for structural entries in the Brookhaven Protein Data Bank produced 719 protein chains with matching mRNA sequence, amino acid sequence, and secondary structure assignment. By neural network analysis, we found strong signals in mRNA sequence regions surrounding helices and sheets. These signals do not originate from the clustering of rare codons, but from the similarity of codons coding for very abundant amino acid residues at the N- and C-termini of helices and sheets. No correlation between the positioning of rare codons and the location of structural units was found. The mRNA signals were also compared with conserved nucleotide features of 16S-like ribosomal RNA sequences and related to mechanisms for maintaining the correct reading frame by the ribosome.

Amino Acid Sequence↗

Cleavage site analysis in picornaviral polyproteins: discovering cellular targets by neural networks.

Picornaviral proteinases are responsible for maturation cleavages of the viral polyprotein, but also catalyze the degradation of cellular targets. Using graphical visualization techniques and neural network algorithms, we have investigated the sequence specificity of the two proteinases 2Apro and 3Cpro. The cleavage of VP0 (giving rise to VP2 and VP4), which is carried out by a so-far unknown proteinase, was also examined. In combination with a novel surface exposure prediction algorithm, our neural network approach successfully distinguishes known cleavage sites from noncleavage sites and yields a more consistent definition of features common to these sites. The method is able to predict experimentally determined cleavage sites in cellular proteins. We present a list of mammalian and other proteins that are predicted to be possible targets for the viral proteinases. Whether these proteins are indeed cleaved awaits experimental verification. Additionally, we report several errors detected in the protein databases. A computer server for prediction of cleavage sites by picornaviral proteinases is publicly available at the e-mail address NetPicoRNA@cbs.dtu.dk or via WWW at http:@www.cbs.dtu.dk/services/NetPicoRNA/.

Amino Acid Sequence↗

Relationship between protein structure and geometrical constraints.

We evaluate to what extent the structure of proteins can be deduced from incomplete knowledge of disulfide bridges, surface assignments, secondary structure assignments, and additional distance constraints. A cost function taking such constraints into account was used to obtain protein structures using a simple minimization algorithm. For small proteins, the approximate structure could be obtained using one additional distance constraint for each amino acid in the protein. We also studied the effect of using predicted secondary structure and surface assignments. The constraints used in this approach typically may be obtained from low-resolution experimental data. When using a cost function based on distances, half of the resulting structures will be mirrored, because the resulting structure and its mirror image will have the same cost. The secondary structure assignments were therefore divided into chirality constraints and distance constraints. Here we report that the problem of mirrored structures, in some cases, can be solved by using a chirality term in the cost function.

Protein Structure, Secondary↗

Distance distributions in proteins: a six-parameter representation.

We present a statistical analysis of protein structures based on interatomic C alpha distances. The overall distance distributions reflect in detail the contents of sequence-specific substructures maintained by local interactions (such as alpha-helixes) and longer range interactions (such as disulfide bridges and beta-sheets). We also show that a volume scaling of the distances makes distance distributions for protein chains of different length superimposable. Distance distributions were also calculated specifically for amino acids separated by a given number of residues. Specific features in these distributions are visible for sequence separations of up to 20 amino acid residues. A simple representation, which preserves most of the information in the distance distributions, was obtained using six parameters only. The parameters give rise to canonical distance intervals and when predicting coarse-grained distance constraints by methods such as data-driven artificial neural networks, these should preferably be selected from these intervals. We discuss the use of the six parameters for determining or reconstructing 3-D protein structures.

Amino Acid Sequence↗

Characterization of prokaryotic and eukaryotic promoters using hidden Markov models.

In this paper we utilize hidden Markov models (HMMs) and information theory to analyze prokaryotic and eukaryotic promoters. We perform this analysis with special emphasis on the fact that promoters are divided into a number of different classes, depending on which polymerase-associated factors that bind to them. We find that HMMs trained on such subclasses of Escherichia coli promoters (specifically, the so-called sigma 70 and sigma 54 classes) give an excellent classification of unknown promoters with respect to sigma-class. HMMs trained on eukaryotic sequences from human genes also model nicely all the essential well known signals, in addition to a potentially new signal upstream of the TATA-box. We furthermore employ a novel technique for automatically discovering different classes in the input data (the promoters) using a system of self-organizing parallel HMMs. These self-organizing HMMs have at the same time the ability to find clusters and the ability to model the sequential structure in the input data. This is highly relevant in situations where the variance in the data is high, as is the case for the subclass structure in for example promoter sequences.

Escherichia coli↗

Analysis of eukaryotic promoter sequences reveals a systematically occurring CT-signal.

A general data study of eukaryotic promoter sequences from widely different species is presented. Mammalian promoters with known transcription initiation sites represented the largest subclass of the data, and for this group neural network algorithms were trained to predict the location of the initiation site in a test set. The prediction accuracy of this local method was higher than what could be expected from the known non-local structure of eukaryotic promoters. Subsequent analysis revealed, besides the consensus of the two known important subregions: the TATA-box TATAAA and the Cap-signal CA, a CT-signal positioned on the average seven nucleotides downstream of the transcription initiation site. The consensus of the CT-signal is CTNCNG. The details of this core promoter element were disclosed using multiple alignment and have earlier only been described in a few isolated examples.

Algorithms↗

Periodic sequence patterns in human exons.

We analyse the sequential structure of human exons and their flanking introns by hidden Markov models. Together, models of donor site regions, acceptor site regions and flanked internal exons, show that exons--besides the reading frame--hold a specific periodic pattern. The pattern, which has the consensus: non-T(A/T)G and a minimal periodicity of roughly 10 nucleotides, is not a consequence of the nucleotide statistics in the three codon positions, nor of the well known nucleosome positioning signal. We discuss the relation between the pattern and other known sequence elements responsible for the intrinsic bending or curvature of DNA.

Base Sequence↗

Neural network model of the genetic code is strongly correlated to the GES scale of amino acid transfer free energies.

A neural network trained to classify the 61 nucleotide triplets of the genetic code into 20 amino acid categories develops in its internal representation a pattern matching the relative cost of transferring amino acids with satisfied backbone hydrogen bonds from water to an environment of dielectric constant of roughly 2.0. Such environments are typically found in lipid membranes or in the interior of proteins. In learning the mapping between the codons and the categories, the network groups the amino acids according to the scale of transfer free energies developed by Engelman, Goldman and Steitz. Several other scales based on internal preference statistics also agree reasonably well with the network grouping. The network is able to relate the structure of the genetic code to quantifications of amino acid hydrophobicity-hydrophilicity more systematically than the numerous attempts made earlier. Due to its inherent non-linearity, the code is also shown to impose decisive constraints on algorithmic analysis of the protein coding potential of DNA.

Amino Acid Sequence↗

Protein structures from distance inequalities.

A computer method for folding protein backbones from distance inequalities is presented. It involves an algorithm that uses a novel approach for handling inequalities through the minimization of a continuous energy function. Tests of the folding algorithm have been carried out on a small protein, the 6PTI (bovine pancreatic trypsin inhibitor) with 56 amino acid residues, and on a medium-size protein, the 1TRM (rat trypsin) with 223 amino acid residues. Reconstructions based on a real-valued distance matrix led to folded three-dimensional structures with root-mean-square values of 0.04 A when compared with the crystallographic data. The obtained root-mean-square measures were of the order of 1 A, when distance inequalities were used for the reconstruction. Subsequently, the folding approach has been applied to distance inequalities predicted by neural network techniques that use the amino acid sequence as the only input. The inaccuracy in the inequalities predicted by the neural network was the reason for the root-mean-square value of 5.2 A. An error analysis of the method for reconstruction was performed and showed that no more than 3% inaccurate distance inequalities could be corrected for. Finally, a simple technique for root-mean-square comparisons of different protein structures is discussed.

Algorithms↗

G+C-rich tract in 5' end of human introns.

Analysis of an artificial neural network trained to classify DNA as coding or non-coding revealed compositional differences between sequence parts translated into protein and those that were not. The 5' end of human introns was found to have a base composition that was non-random to an extent matching the non-randomness in the 3' end that contains the polypyrimidine tract. The prevailing nucleotides in the initial 50 nucleotides of human introns are guanine and cytosine, the trinucleotide GGG was found to occur almost four times as frequently as it would in sequences with a uniform distribution of the nucleotides. The initial part of terminal exons and their associated terminal introns were shown to have a very special base composition deviating strongly from the normal picture in other exons and introns.

Base Composition↗

Multiple alignment using simulated annealing: branch point definition in human mRNA splicing.

A method for the simultaneous alignment of a very large number of sequences using simulated annealing is presented. The total running time of the algorithm does not depend explicitly on the number of sequences treated. The method has been used for the simultaneous alignment of 1462 human intron sequences upstream of the intron-exon boundary. The consensus sequence of the aligned set together with a calculation of the Shannon information clearly shows that several sequence motives are conserved: (i) a previously undetected guanosine rich region, (ii) the branch point and (iii) the polypyrimidine tract. The nucleotide frequencies at each position of the branch point consensus sequence qualitatively reproduce the frequencies of the experimentally determined branch points.

Algorithms↗

Prediction of human mRNA donor and acceptor sites from the DNA sequence.

Artificial neural networks have been applied to the prediction of splice site location in human pre-mRNA. A joint prediction scheme where prediction of transition regions between introns and exons regulates a cutoff level for splice site assignment was able to predict splice site locations with confidence levels far better than previously reported in the literature. The problem of predicting donor and acceptor sites in human genes is hampered by the presence of numerous amounts of false positives: here, the distribution of these false splice sites is examined and linked to a possible scenario for the splicing mechanism in vivo. When the presented method detects 95% of the true donor and acceptor sites, it makes less than 0.1% false donor site assignments and less than 0.4% false acceptor site assignments. For the large data set used in this study, this means that on average there are one and a half false donor sites per true donor site and six false acceptor sites per true acceptor site. With the joint assignment method, more than a fifth of the true donor sites and around one fourth of the true acceptor sites could be detected without accompaniment of any false positive predictions. Highly confident splice sites could not be isolated with a widely used weight matrix method or by separate splice site networks. A complementary relation between the confidence levels of the coding/non-coding and the separate splice site networks was observed, with many weak splice sites having sharp transitions in the coding/non-coding signal and many stronger splice sites having more ill-defined transitions between coding and non-coding.

Base Sequence↗

Neural network detects errors in the assignment of mRNA splice sites.

The use of databanks in genetic research assumes reliability of the information they contain. Currently, error-detection in the manually or electronically entered data contained in the nucleotide sequence databanks at EMBL, Heidelberg and GenBank at Los Alamos is limited. We have used a subset of sequences from these databanks to train neural networks to recognize pre-mRNA splicing signals in human genes. During the training on 33 human genes from the EMBL databank seven genes appeared to disturb the learning process. Subsequent investigation revealed discrepancies from the original published papers, for three genes. In four genes, we found wrongly assigned splicing frames of introns. We believe this to be a reflection of the fact that splicing frames cannot always be unambiguously assigned on the basis of experimental data. Thus incorrect assignment appear both due to mere typographical misprints as well as erroneous interpretation of experiments. Training on 241 human sequences from GenBank revealed nine new errors. We propose that such errors could be detected by computer algorithms designed to check the consistency of data prior to their incorporation in databanks.

Algorithms↗

Analysis of the secondary structure of the human immunodeficiency virus (HIV) proteins p17, gp120, and gp41 by computer modeling based on neural network methods.

A neural network computer program, trained to predict secondary structure of proteins by exposing it to matching sets of primary and secondary structures from a database, was used to analyze the human immunodeficiency virus (HIV) proteins p17, gp120, and gp41 from their amino acid sequences. The results are compared to those obtained by the Chou-Fasman analysis. Two alpha-helical sequences corresponding to the putative fusigenic domain and to the transmembrane domain of gp41 could be predicted, as well as a possible binding site between p17 and gp41. On the basis of the secondary structure predictions, a three-dimensional model of p17 was constructed. This model was found to represent a stable conformation by an analysis using an energy-minimization program. The model predicts that p17 is attached to the membrane only by the acylated N-terminus, in analogy with the N-terminus of the gag protein of other retroviruses and also with the src oncogene protein p60src. The intracellular C-terminal part of gp41 may act as a receptor by electrostatic interaction with p17.

Algorithms↗

Protein secondary structure and homology by neural networks. The alpha-helices in rhodopsin.

Neural networks provide a basis for semiempirical studies of pattern matching between the primary and secondary structures of proteins. Networks of the perceptron class have been trained to classify the amino-acid residues into two categories for each of three types of secondary feature: alpha-helix or not, beta-sheet or not, and random coil or not. The explicit prediction for the helices in rhodopsin is compared with both electron microscopy results and those of the Chou-Fasman method. A new measure of homology between proteins is provided by the network approach, which thereby leads to quantification of the differences between the primary structures of proteins.

Amino Acid Sequence↗

Improving the odds in discriminating "drug-like" from "non drug-like" compounds.

We have used a feed-forward neural network technique to classify chemical compounds into potentially "drug-like" and "non drug-like" candidates. The neural network was trained to distinguish between a set of "drug-like" and "non drug-like" chemical compounds taken from the MACCS-II Drug Data Report (MDDR) and the Available Chemicals Directory (ACD). The 2D atom types (of the full atomic representation) were assigned and applied as descriptors to encode numerically each compound. There are four main conclusions: First the method performs well, correctly assigning 88% of the compounds in both MDDR and ACD. Improved discrimination was achieved by a more critical selection of training sets. Second, the method gives much better prediction performance than the widely used "Rule of Five", which accepts as many as 74% of the ACD compounds but only 66% of those in MDDR, resulting in a correlation coefficient which is effectively zero, compared to a value of 0.63 for the neural network prediction. Third, based on a standard Tanimoto similarity search the selection of drug-like compounds in the evaluation set is not biased toward compounds similar to those in the training set. Fourth, the trained neural network was applied to evaluate the drug-likeness of 136 GABA uptake inhibitors with impressive results. The implications of applying a neural network to characterize chemical compounds are discussed.

Algorithms↗