Search PubMed⌕ Search

Biomedical subjects

P Baldi

Publications and source records attributed to P Baldi.

At least 37 records · Page 2Linked to original sources

The biology of eukaryotic promoter prediction--a review.

Computational prediction of eukaryotic promoters from the nucleotide sequence is one of the most attractive problems in sequence analysis today, but it is also a very difficult one. Thus, current methods predict in the order of one promoter per kilobase in human DNA, while the average distance between functional promoters has been estimated to be in the range of 30-40 kilobases. Although it is conceivable that some of these predicted promoters correspond to cryptic initiation sites that are used in vivo, it is likely that most are false positives. This suggests that it is important to carefully reconsider the biological data that forms the basis of current algorithms, and we here present a review of data that may be useful in this regard. The review covers the following topics: (1) basal transcription and core promoters, (2) activated transcription and transcription factor binding sites, (3) CpG islands and DNA methylation, (4) chromosomal structure and nucleosome modification, and (5) chromosomal domains and domain boundaries. We discuss the possible lessons that may be learned, especially with respect to the wealth of information about epigenetic regulation of transcription that has been appearing in recent years.

Chromosomes↗

Nackt (nkt), a new hair loss mutation of the mouse with associated CD4 deficiency.

A spontaneous recessive mutation named nackt (symbol: nkt) affecting hair growth and T-cell development was discovered in a moderately inbred stock of mice. Skin lesions were characterized by sparse rough coat, bare patches around the eyes and neck, and a scratching behavior throughout life. Fluorescence-activated cell sorter analysis indicated a deficiency in the CD4(+) 8(-) T-cell subset in the thymus and a marked decrease in CD4(+) T cells in peripheral lymphoid organs. Linkage analysis using a set of molecular markers and an F2 intersubspecific cross indicated that the mutation maps to the central region of mouse chromosome 13, in a region homologous to human chromosome 5q22-q35.

Alopecia↗

High expression level of a gene coding for a chloroplastic amino acid selective channel protein is correlated to cold acclimation in cereals.

A cold-regulated gene (cor tmc-ap3) coding for a putative chloroplastic amino acid selective channel protein was isolated from cold-treated barley leaves combining the differential display and the 5'-RACE techniques. Cor tmc-ap3 is expressed at low level under normal growing temperature, and its expression is strongly enhanced after cold treatment. A positive correlation between the expression of cor tmc-ap3 and frost tolerance was found both among barley cultivars and among cereal species. The COR TMC-AP3 protein was expressed in vitro, purified and used to raise a polyclonal antibody. Western analysis showed that the cor tmc-ap3 gene product is localized to the chloroplastic outer envelope fraction, supporting its putative function. The frost-resistant winter cultivar Onice accumulated COR TMC-AP3 more rapidly and at a higher level than the frost-susceptible spring cultivar Gitane. After 28 days of cold acclimation the winter cultivar had about 2-fold more protein than the spring genotype. All these results suggest that an increased amount of a chloroplastic amino acid selective channel protein could be required for cold acclimation in cereals. Hypotheses about the role of COR TMC-AP3 during the hardening process are discussed.

Acclimatization↗

Structural basis for triplet repeat disorders: a computational analysis.

MOTIVATION: Over a dozen major degenerative disorders, including myotonic distrophy, Huntington's disease and fragile X syndrome, result from unstable expansions of particular trinucleotides. Remarkably, only some of all the possible triplets, namely CAG/CTG, CGG/CCG and GAA/TTC, have been associated with the known pathological expansions. This raises some basic questions at the DNA level. Why do particular triplets seem to be singled out? What is the mechanism for their expansion and how does it depend on the triplet itself? Could other triplets or longer repeats be involved in other diseases? RESULTS: Using several different computational models of DNA structure, we show that the triplets involved in the pathological repeats generally fall into extreme classes. Thus, CAG/CTG repeats are particularly flexible, whereas GCC, CGG and GAA repeats appear to display both flexible and rigid (but curved) characteristics depending on the method of analysis. The fact that (1) trinucleotide repeats often become increasingly unstable when they exceed a length of approximately 50 repeats, and (2) repeated 12-mers display a similar increase in instability above 13 repeats, together suggest that approximately 150 bp is a general threshold length for repeat instability. Since this is about the length of DNA wrapped up in a single nucleosome core particle, we speculate that chromatin structure may play an important role in the expansion mechanism. We furthermore suggest that expansion of a dodecamer repeat, which we predict to have very high flexibility, may play a role in the pathogenesis of the neurodegenerative disorder multiple system atrophy (MSA). CONTACT: pfbaldi@ics.uci.edu, yves@netid.com, brunak@cbs.dtu.dk, gorm@cbs.dtu.dk.

Anticipation, Genetic↗

Exploiting the past and the future in protein secondary structure prediction.

MOTIVATION: Predicting the secondary structure of a protein (alpha-helix, beta-sheet, coil) is an important step towards elucidating its three-dimensional structure, as well as its function. Presently, the best predictors are based on machine learning approaches, in particular neural network architectures with a fixed, and relatively short, input window of amino acids, centered at the prediction site. Although a fixed small window avoids overfitting problems, it does not permit capturing variable long-rang information. RESULTS: We introduce a family of novel architectures which can learn to make predictions based on variable ranges of dependencies. These architectures extend recurrent neural networks, introducing non-causal bidirectional dynamics to capture both upstream and downstream information. The prediction algorithm is completed by the use of mixtures of estimators that leverage evolutionary information, expressed in terms of multiple alignments, both at the input and output levels. While our system currently achieves an overall performance close to 76% correct prediction--at least comparable to the best existing systems--the main emphasis here is on the development of new algorithmic ideas. AVAILABILITY: The executable program for predicting protein secondary structure is available from the authors free of charge. CONTACT: pfbaldi@ics.uci.edu, gpollast@ics.uci.edu, brunak@cbs.dtu.dk, paolo@dsi.unifi.it.

Algorithms↗

DNA structure in human RNA polymerase II promoters.

The fact that DNA three-dimensional structure is important for transcriptional regulation begs the question of whether eukaryotic promoters contain general structural features independently of what genes they control. We present an analysis of a large set of human RNA polymerase II promoters with a very low level of sequence similarity. The sequences, which include both TATA-containing and TATA-less promoters, are aligned by hidden Markov models. Using three different models of sequence-derived DNA bendability, the aligned promoters display a common structural profile with bendability being low in a region upstream of the transcriptional start point and significantly higher downstream. Investigation of the sequence composition in the two regions shows that the bendability profile originates from the sequential structure of the DNA, rather than the general nucleotide composition. Several trinucleotides known to have high propensity for major groove compression are found much more frequently in the regions downstream of the transcriptional start point, while the upstream regions contain more low-bendability triplets. Within the region downstream of the start point, we observe a periodic pattern in sequence and bendability, which is in phase with the DNA helical pitch. The periodic bendability profile shows bending peaks roughly at every 10 bp with stronger bending at 20 bp intervals. These observations suggest that DNA in the region downstream of the transcriptional start point is able to wrap around protein in a manner reminiscent of DNA in a nucleosome. This notion is further supported by the finding that the periodic bendability is caused mainly by the complementary triplet pairs CAG/CTG and GGC/GCC, which previously have been found to correlate with nucleosome positioning. We present models where the high-bendability regions position nucleosomes at the downstream end of the transcriptional start point, and consider the possibility of interaction between histone-like TAFs and this area. We also propose the use of this structural signature in computational promoter-finding algorithms.

Algorithms↗

On the use of Bayesian methods for evaluating compartmental neural models.

Computational modeling is being used increasingly in neuroscience. In deriving such models, inference issues such as model selection, model complexity, and model comparison must be addressed constantly. In this article we present briefly the Bayesian approach to inference. Under a simple set of commonsense axioms, there exists essentially a unique way of reasoning under uncertainty by assigning a degree of confidence to any hypothesis or model, given the available data and prior information. Such degrees of confidence must obey all the rules governing probabilities and can be updated accordingly as more data becomes available. While the Bayesian methodology can be applied to any type of model, as an example we outline its use for an important, and increasingly standard, class of models in computational neuroscience--compartmental models of single neurons. Inference issues are particularly relevant for these models: their parameter spaces are typically very large, neurophysiological and neuroanatomical data are still sparse, and probabilistic aspects are often ignored. As a tutorial, we demonstrate the Bayesian approach on a class of one-compartment models with varying numbers of conductances. We then apply Bayesian methods on a compartmental model of a real neuron to determine the optimal amount of noise to add to the model to give it a level of spike time variability comparable to that found in the real cell.

Action Potentials↗

Computational applications of DNA structural scales.

We study from a computational standpoint several different physical scales associated with structural features of DNA sequences, including dinucleotide scales such as base stacking energy and propeller twist, and trinucleotide scales such as bendability and nucleosome positioning. We show that these scales provide an alternative or complementary compact representation of DNA sequences. As an example we construct a strand invariant representation of DNA sequences. The scales can also be used to analyze and discover new DNA structural patterns, especially in combinations with hidden Markov models (HMMs). The scales are applied to HMMs of human promoter sequences revealing a number of significant differences between regions upstream and downstream of the transcriptional start point. Finally we show, with some qualifications, that such scales are by and large independent, and therefore complement each other.

Artificial Intelligence↗

[In vitro evaluation of the cytotoxicity of composite resin in the presence or absence of smear layer].

BACKGROUND: The aim of the present research was to investigate the cytotoxicity of the composite resins applied on dentine samples with different smear layer removal, by using cell cultures and LDH determination. MATERIALS AND METHODS: Ninety-eight caries-free third molars recently extracted were used. Transversal sections, 500 mu thick, were obtained. A simulated pulpar chamber was constructed allowing to give lodging in the inferior side to the cell culture and in the superior side to the dentine section with the composite resin. LDH activity was then determined by using a spectrophotometer. BACKGROUND: The results show that the composite resins are surely cytotoxic if directly applied on the dentine. The smear layer is able to reduce the transdentinal diffusion of composite resin toxicity. On the basis of the data obtained it is suggested that in vivo, being necessary to eliminate the smear layer due to its bacterial contents, it is possible in the deep cavities, to partially remove with EDTA maintaining the smear plugs after their disinfection. Nevertheless EDTA application should not exceed 30 seconds.

Cell Culture Techniques↗

Naturally occurring nucleosome positioning signals in human exons and introns.

We describe the structural implications of a periodic pattern found in human exons and introns by hidden Markov models. We show that exons (besides the reading frame) have a specific sequential structure in the form of a pattern with triplet consensus non-T(A/T)G, and a minimal periodicity of roughly ten nucleotides. The periodic pattern is also present in intron sequences, although the strength per nucleotide is weaker. Using two independent profile methods based on triplet bendability parameters from DNase I experiments and nucleosome positioning data, we show that the pattern in multiple alignments of internal exon and intron sequences corresponds to a periodic "in phase" bending potential towards the major groove of the DNA. The nucleosome positioning data show that the consensus triplets (and their complements) have a preference for locations on a bent double helix where the major groove faces inward and is compressed. The in-phase triplets are located adjacent to GCC/GGC triplets known to have the strongest bias in their positioning on the nuclesome. Analysis of mRNA sequences encoding proteins with known tertiary structure exclude the possibility that the pattern is a consequence of the previously well-known periodicity caused by the encoding of alpha-helices in proteins. Finally, we discuss the relation between the bending potential of coding and non-coding regions and its impact on the translational positioning of nucleosomes and the recognition of genes by the transcriptional machinery.

Base Sequence↗

Hybrid modeling, HMM/NN architectures, and protein applications.

We describe a hybrid modeling approach where the parameters of a mode are calculated and modulated by another model, typically a neural network (NN), to avoid both overfitting and underfitting. We develop the approach for the case of Hidden Markov Models (HMMs), by deriving a class of hybrid HMM/NN architectures. These architectures can be trained with unified algorithms that blend HMM dynamic programming with NN backpropagation. In the case of complex data, mixtures of HMMs or modulated HMMs must be used. NNs can then be applied both to the parameters of each single HMM, and to the switching or modulatation of the models, as a function of input or context. Hybrid HMM/NN architectures provide a flexible NN parameterization for the control of model structure and complexity. At the same time, they can capture distributions that, in practice, are inaccessible to single HMMs. The HMM/NN hybrid approach is tested, in its simplest form, by constructing a model of the immunoglobulin protein family. A hybrid model is trained, and a multiple alignment derived, with less than a fourth of the number of parameters used with previous single HMMs.

Algorithms↗

Characterization of prokaryotic and eukaryotic promoters using hidden Markov models.

In this paper we utilize hidden Markov models (HMMs) and information theory to analyze prokaryotic and eukaryotic promoters. We perform this analysis with special emphasis on the fact that promoters are divided into a number of different classes, depending on which polymerase-associated factors that bind to them. We find that HMMs trained on such subclasses of Escherichia coli promoters (specifically, the so-called sigma 70 and sigma 54 classes) give an excellent classification of unknown promoters with respect to sigma-class. HMMs trained on eukaryotic sequences from human genes also model nicely all the essential well known signals, in addition to a potentially new signal upstream of the TATA-box. We furthermore employ a novel technique for automatically discovering different classes in the input data (the promoters) using a system of self-organizing parallel HMMs. These self-organizing HMMs have at the same time the ability to find clusters and the ability to model the sequential structure in the input data. This is highly relevant in situations where the variance in the data is high, as is the case for the subclass structure in for example promoter sequences.

Escherichia coli↗

Substitution matrices and hidden Markov models.

Hidden Markov models (HMMs) provide a general framework for expressing primary sequence consensus. HMMs can effectively be used to model and align protein families, and to search data bases. HMMs, however, have a large number of parameters. When only few sequences are available for model fitting, additional prior information must be incorporated into the models. We derive a simple algorithm that directly incorporates prior information provided by substitution matrices into the HMM learning procedure.

Algorithms↗

Periodic sequence patterns in human exons.

We analyse the sequential structure of human exons and their flanking introns by hidden Markov models. Together, models of donor site regions, acceptor site regions and flanked internal exons, show that exons--besides the reading frame--hold a specific periodic pattern. The pattern, which has the consensus: non-T(A/T)G and a minimal periodicity of roughly 10 nucleotides, is not a consequence of the nucleotide statistics in the three codon positions, nor of the well known nucleosome positioning signal. We discuss the relation between the pattern and other known sequence elements responsible for the intrinsic bending or curvature of DNA.

Base Sequence↗

Protein modeling with hybrid Hidden Markov Model/neural network architectures.

Hidden Markov Models (HMMs) are useful in a number of tasks in computational molecular biology, and in particular to model and align protein families. We argue that HMMs are somewhat optimal within a certain modeling hierarchy. Single first order HMMs, however, have two potential limitations: a large number of unstructured parameters, and a built-in inability to deal with long-range dependencies. Hybrid HMM/Neural Network (NN) architectures attempt to overcome these limitations. In hybrid HMM/NN, the HMM parameters are computed by a NN. This provides a reparametrization that allows for flexible control of model complexity, and incorporation of constraints. The approach is tested on the immunoglobulin family. A hybrid model is trained, and a multiple alignment derived, with less than a fourth of the number of parameters used with previous single HMMs. To capture dependencies, however, one must resort to a larger hybrid model class, where the data is modeled by multiple HMMs. The parameters of the HMMs, and their modulation as a function of input or context, is again calculated by a NN.

Amino Acid Sequence↗

Hidden Markov models of biological primary sequence information.

Hidden Markov model (HMM) techniques are used to model families of biological sequences. A smooth and convergent algorithm is introduced to iteratively adapt the transition and emission parameters of the models from the examples in a given family. The HMM approach is applied to three protein families: globins, immunoglobulins, and kinases. In all cases, the models derived capture the important statistical characteristics of the family and can be used for a number of tasks, including multiple alignments, motif detection, and classification. For K sequences of average length N, this approach yields an effective multiple-alignment algorithm which requires O(KN2) operations, linear in the number of sequences.

Algorithms↗

Hidden Markov Models of the G-protein-coupled receptor family.

Hidden Markov Model techniques are used to derive a new model of the G-protein-coupled receptor family. The transition and emission parameters of the model are adjusted using a training set comprising 142 sequences. The resulting model is shown to perform well on a number of tasks, including multiple alignments, discrimination, large data base searches, classification, and fragment detection. General analytical results on the expectation and standard deviation of the likelihood of random sequences are also presented.

Algorithms↗