Search PubMed⌕ Search

Biomedical subjects

Julio Collado-Vides

Publications and source records attributed to Julio Collado-Vides.

5 recordsLinked to original sources

Evaluation of thresholds for the detection of binding sites for regulatory proteins in Escherichia coli K12 DNA.

BACKGROUND: Sites in DNA that bind regulatory proteins can be detected computationally in various ways. Pattern discovery methods analyze collections of genes suspected to be co-regulated on the evidence, for example, of clustering of transcriptome data. Pattern searching methods use sequences with known binding sites to find other genes regulated by a given protein. Such computational methods are important strategies in the discovery and elaboration of regulatory networks and can provide the experimental biologist with a precise prediction of a binding site or identify a gene as a member of a set of co-regulated genes (a regulon). As more variations on such methods are published, however, thorough evaluation is necessary, as performance may differ depending on the conditions of use. Detailed evaluation also helps to improve and understand the behavior of the different methods and computational strategies. RESULTS: We used a collection of 86 regulons from Escherichia coli as datasets to evaluate two methods for pattern discovery and pattern searching: dyad analysis/dyad sweeping using the program Dyad-analysis, and multiple alignment using the programs Consensus/Patser. Clearly defined statistical parameters are used to evaluate the two methods in different situations. We placed particular emphasis on minimizing the rate of false positives. CONCLUSIONS: As a general rule, sensors obtained from experimentally reported binding sites in DNA frequently locate true sites as the highest-scoring sequences within a given upstream region, especially using Consensus/Patser. Pattern discovery is still an unsolved problem, although in the cases where Dyad-analysis finds significant dyads (around 50%), these frequently correspond to true binding sites. With more robust methods, regulatory predictions could help identify the function of unknown genes.

Bacterial Proteins↗

The EcoCyc Database.

EcoCyc is an organism-specific pathway/genome database that describes the metabolic and signal-transduction pathways of Escherichia coli, its enzymes, its transport proteins and its mechanisms of transcriptional control of gene expression. EcoCyc is queried using the Pathway Tools graphical user interface, which provides a wide variety of query operations and visualization tools. EcoCyc is available at http://ecocyc.org/.

Database Management Systems↗

Analysis of the cellular functions of Escherichia coli operons and their conservation in Bacillus subtilis.

The common assumption of operons as composed of genes that cooperate in a biological process is confirmed here by showing that Escherichia coli operons tend to be composed of genes that belong to the same general class of cellular function. Furthermore, the comparison between the genomic organization of E. coli and that of Bacillus subtilis shows that the genes that are homologous to genes that belong to experimentally characterized E. coli operons tend to cluster in neighboring regions of the genome. This tendency is greater for the subset of E. coli operons whose genes belong to a single functional class. These observations indicate strong evolutionary pressure that, translated into functional constraints, leads to the inclusion of many essential functions in conserved operons and clusters in these two distant species.

Bacillus subtilis↗

A powerful non-homology method for the prediction of operons in prokaryotes.

MOTIVATION: The prediction of the transcription unit organization of genomes is an important clue in the inference of functional relationships of genes, the interpretation and evaluation of transcriptome experiments, and the overall inference of the regulatory networks governing the expression of genes in response to the environment. Though several methods have been devised to predict operons, most need a high characterization of the genome analysed. Log-likelihoods derived from inter-genic distance distributions work surprisingly well to predict operons in Escherichia coli and are available for any genome as soon as the gene sets are predicted. RESULTS: Here we provide evidence that the very same method is applicable to any prokaryotic genome. First, the method has the same efficiency when evaluated using a collection of experimentally known operons of Bacillus subtilis. Second, operons among most if not all prokaryotes seem to have the same tendencies to keep short distances between their genes, the most frequent distances being the overlaps of four and one base pairs. The universality of this structural feature allows us to predict the organization of transcription units in all prokaryotes. Third, predicted operons contain a higher proportion of genes with related phylogenetic profiles and conservation of adjacency than predicted borders of transcription units.

Algorithms↗

Operon conservation from the point of view of Escherichia coli, and inference of functional interdependence of gene products from genome context.

We have previously demonstrated that genes within experimentally characterized operons of Escherichia coli are conserved together in other genomes more frequently than genes at the borders of transcription units. Here we expand the analyses and show that, as the phylogenetic distance of the genomes compared increases, the genes remaining together must belong to genes associated into operons in other prokaryotes regardless of the operon organization of the corresponding orthologous gene pair of E. coli. At the same time, we show that the observed tendencies of genes within operons to keep very short inter-genic distances in E. coli, is the same in any other prokaryote whose genome is currently available. We also show the relationship between our analyses of conservation and the inference of functional relationships from genomic context.

Computational Biology↗