Search PubMed⌕ Search

Biomedical subjects

Graziano Pesole

Publications and source records attributed to Graziano Pesole.

At least 19 recordsLinked to original sources

Establishing the ELIXIR Microbiome Community.

Microbiome research has grown substantially over the past decade in terms of the range of biomes sampled, identified taxa, and the volume of data derived from the samples. In particular, experimental approaches such as metagenomics, metabarcoding, metatranscriptomics and metaproteomics have provided profound insights into the vast, hitherto unknown, microbial biodiversity. The ELIXIR Marine Metagenomics Community, initiated amongst researchers focusing on marine microbiomes, has concentrated on promoting standards around microbiome-derived sequence analysis, as well as understanding the gaps in methods and reference databases, and identifying solutions to the computational overheads of performing such analyses. Nevertheless, the methods used and the challenges faced are not confined to marine microbiome studies, but are broadly applicable to other biomes. Thus, expanding this Marine Metagenomics Community to a more inclusive ELIXIR Microbiome Community will enable it to encompass a broader range of biomes and link expertise across 'omics technologies. Furthermore, engaging with a large number of researchers will improve the efficiency and sustainability of bioinformatics infrastructure and resources for microbiome research (standards, data, tools, workflows, training), which will enable a deeper understanding of the function and taxonomic composition of the different microbial communities.

Computational Biology↗

DG-CST (Disease Gene Conserved Sequence Tags), a database of human-mouse conserved elements associated to disease genes.

The identification and study of evolutionarily conserved genomic sequences that surround disease-related genes is a valuable tool to gain insight into the functional role of these genes and to better elucidate the pathogenetic mechanisms of disease. We created the DG-CST (Disease Gene Conserved Sequence Tags) database for the identification and detailed annotation of human-mouse conserved genomic sequences that are localized within or in the vicinity of human disease-related genes. CSTs are defined as sequences that show at least 70% identity between human and mouse over a length of at least 100 bp. The database contains CST data relative to over 1088 genes responsible for monogenetic human genetic diseases or involved in the susceptibility to multifactorial/polygenic diseases. DG-CST is accessible via the internet at http://dgcst.ceinge.unina.it/ and may be searched using both simple and complex queries. A graphic browser allows direct visualization of the CSTs and related annotations within the context of the relative gene and its transcripts.

Animals↗

UTRdb and UTRsite: a collection of sequences and regulatory motifs of the untranslated regions of eukaryotic mRNAs.

The 5' and 3' untranslated regions of eukaryotic mRNAs play crucial roles in the post-transcriptional regulation of gene expression through the modulation of nucleo-cytoplasmic mRNA transport, translation efficiency, subcellular localization and message stability. UTRdb is a curated database of 5' and 3' untranslated sequences of eukaryotic mRNAs, derived from several sources of primary data. Experimentally validated functional motifs are annotated (and also collated as the UTRsite database) and cross-links to genomic and protein data are provided. The integration of UTRdb with genomic and protein data has allowed the implementation of a powerful retrieval resource for the selection and extraction of UTR subsets based on their genomic coordinates and/or features of the protein encoded by the relevant mRNA (e.g. GO term, PFAM domain, etc.). All internet resources implemented for retrieval and functional analysis of 5' and 3' untranslated regions of eukaryotic mRNAs are accessible at http://www.ba.itb.cnr.it/UTR/.

3' Untranslated Regions↗

Energy biogenesis: one key for coordinating two genomes.

In metazoan organisms, energy production is the only example of a process that is under dual genetic control: nuclear and mitochondrial. We used a genomic approach to examine how energy genes of both the nuclear and mitochondrial genomes are coordinated, and discovered a novel genetic regulatory circuit in Drosophila melanogaster that is surprisingly simple and parsimonious. This circuit is based on a single DNA regulatory element and can explain both intra- and inter-genomic coordinated expression of genes involved in energy production, including the full complement of mitochondrial and nuclear oxidative phosphorylation genes, and the genes involved in the Krebs cycle.

Animals↗

Assessing computational tools for the discovery of transcription factor binding sites.

The prediction of regulatory elements is a problem where computational methods offer great hope. Over the past few years, numerous tools have become available for this task. The purpose of the current assessment is twofold: to provide some guidance to users regarding the accuracy of currently available tools in various settings, and to provide a benchmark of data sets for assessing future tools.

Amino Acid Motifs↗

DNAfan: a software tool for automated extraction and analysis of user-defined sequence regions.

SUMMARY: DNAfan (DNA Feature ANalyzer) is a tool combining sequence-filtering and pattern searching. DNAfan automatically extracts user-defined sets of sequence fragments from large sequence sets. Fragments are defined by annotated gene feature keys and co- or non-occurring patterns within the feature or close to it. A gene feature parser and a pattern-based filter tool localizes and extracts the specific subset of sequences. The selected sequence data can subsequently be retrieved for analyses or further processed with DNAfan to find the occurrence of specific patterns or structural motifs. DNAfan is a powerful tool for pattern analysis. Its filter features restricts the pattern search to a well-defined set of sequences, allowing drastic reduction in false positive hits. AVAILABILITY: http://bighost.ba.itb.cnr.it:8080/Framework.

Algorithms↗

Weeder Web: discovery of transcription factor binding sites in a set of sequences from co-regulated genes.

One of the greatest challenges that modern molecular biology is facing is the understanding of the complex mechanisms regulating gene expression. A fundamental step in this process requires the characterization of regulatory motifs playing key roles in the regulation of gene expression at transcriptional and post-transcriptional levels. In particular, transcription is modulated by the interaction of transcription factors with their corresponding binding sites. Weeder Web is a web interface to Weeder, an algorithm for the automatic discovery of conserved motifs in a set of related regulatory DNA sequences. The motifs found are in turn likely to be instances of binding sites for some transcription factor. Other than providing access to the program, the interface has been designed so to make usage of the program itself as simple as possible, and to require very little prior knowledge about the length and the conservation of the motifs to be found. In fact, the interface automatically starts different runs of the program, each one with different parameters, and provides the user with an overall summary of the results as well as some 'advice' on which motifs look more interesting according to their statistical significance and some simple considerations. The web interface is available at the address www.pesolelab.it by following the 'Tools' link.

Algorithms↗

CSTminer: a web tool for the identification of coding and noncoding conserved sequence tags through cross-species genome comparison.

The identification and characterization of genome tracts that are highly conserved across species during evolution may contribute significantly to the functional annotation of whole-genome sequences. Indeed, such sequences are likely to correspond to known or unknown coding exons or regulatory motifs. Here, we present a web server implementing a previously developed algorithm that, by comparing user-submitted genome sequences, is able to identify statistically significant conserved blocks and assess their coding or noncoding nature through the measure of a coding potential score. The web tool, available at http://www.caspur.it/CSTminer/, is dynamically interconnected with the Ensembl genome resources and produces a graphical output showing a map of detected conserved sequences and annotated gene features.

Animals↗

RNAProfile: an algorithm for finding conserved secondary structure motifs in unaligned RNA sequences.

The recent interest sparked due to the discovery of a variety of functions for non-coding RNA molecules has highlighted the need for suitable tools for the analysis and the comparison of RNA sequences. Many trans-acting non-coding RNA genes and cis-acting RNA regulatory elements present motifs, conserved both in structure and sequence, that can be hardly detected by primary sequence analysis alone. We present an algorithm that takes as input a set of unaligned RNA sequences expected to share a common motif, and outputs the regions that are most conserved throughout the sequences, according to a similarity measure that takes into account both the sequence of the regions and the secondary structure they can form according to base-pairing and thermodynamic rules. Only a single parameter is needed as input, which denotes the number of distinct hairpins the motif has to contain. No further constraints on the size, number and position of the single elements comprising the motif are required. The algorithm can be split into two parts: first, it extracts from each input sequence a set of candidate regions whose predicted optimal secondary structure contains the number of hairpins given as input. Then, the regions selected are compared with each other to find the groups of most similar ones, formed by a region taken from each sequence. To avoid exhaustive enumeration of the search space and to reduce the execution time, a greedy heuristic is introduced for this task. We present different experiments, which show that the algorithm is capable of characterizing and discovering known regulatory motifs in mRNA like the iron responsive element (IRE) and selenocysteine insertion sequence (SECIS) stem-loop structures. We also show how it can be applied to corrupted datasets in which a motif does not appear in all the input sequences, as well as to the discovery of more complex motifs in the non-coding RNA.

3' Untranslated Regions↗

WebVar: A resource for the rapid estimation of relative site variability from multiple sequence alignments.

UNLABELLED: WebVar is an online resource that provides estimates of relative site variability from multiple alignments of homologous protein or nucleic acid sequences. WebVar provides a variety of graphic and textual representations of estimates, designed to assist in phylogenetic analysis. AVAILABILITY: The WebVar server is located at http://www.pesolelab.it/Tools/WebVar.html

Algorithms↗

Complete mtDNA of Ciona intestinalis reveals extensive gene rearrangement and the presence of an atp8 and an extra trnM gene in ascidians.

The complete mitochondrial genome (mtDNA) of the model organism Ciona intestinalis (Urochordata, Ascidiacea) has been amplified by long-PCR using specific primers designed on putative mitochondrial transcripts identified from publicly available mitochondrial-like expressed sequence tags. The C. intestinalis mtDNA encodes 39 genes: 2 rRNAs, 13 subunits of the respiratory complexes, including ATPase subunit 8 ( atp8), and 24 tRNAs, including 2 tRNA-Met with anticodons 5'-UAU-3'and 5'-CAU-3', respectively. All genes are transcribed from the same strand. This gene content seems to be a common feature of ascidian mtDNAs, as we have verified the presence of a previously undetected atp8 and of two trnM genes in the two other sequenced ascidian mtDNAs. Extensive gene rearrangement has been found in C. intestinalis with respect not only to the common Vertebrata/Cephalochordata/Hemichordata gene organization but also to other ascidian mtDNAs, including the cogeneric Ciona savignyi. Other features such as the absence of long noncoding regions, the shortness of rRNA genes, the low GC content (21.4%), and the absence of asymmetric base distribution between the two strands suggest that this genome is more similar to those of some protostomes than to deuterostomes.

Amino Acid Sequence↗

In silico representation and discovery of transcription factor binding sites.

Understanding the complex mechanisms governing basic biological processes requires the characterisation of regulatory motifs modulating gene expression at transcriptional and post-transcriptional level. In particular, extent, chronology and cell-specificity of transcription are modulated by the interaction of transcription factors with their corresponding binding sites, mostly located near (or sometimes quite far away from) the transcription start site of the gene. The constantly growing amount of genomic data, complemented by other sources of information such as expression data derived from microarray experiments, has opened new opportunities to researchers in this field. Many different methods have been proposed for the identification of transcription factor binding sites in the regulatory regions of co-expressed genes: unfortunately this is a very challenging problem both from the computational and the biological viewpoint. This paper provides a survey of existing methods proposed for the problem, focusing both on the ideas underlying them and their availability to the scientific community.

Algorithms↗

Phylogenetic analyses: a brief introduction to methods and their application.

Phylogenetic analysis of molecular sequence data plays an increasingly important role in clinical medicine, both in the emerging field of molecular epidemiology and in the rational design of new therapeutic agents. The aims of this review are to introduce some of the methods used to construct phylogenetic trees, to illustrate some of the pitfalls that can introduce artifactual results and to speculate on the long-term importance of this area of computational biology in clinical medicine.

Animals↗

Congruent mammalian trees from mitochondrial and nuclear genes using Bayesian methods.

Analyses of mitochondrial and nuclear gene sequences have often produced different mammalian tree topologies, undermining confidence in the merit of molecular approaches with respect to "traditional" morphological classification. The recent sequencing of the complete mitochondrial genomes of two additional rodents (Spalax judaei and Jaculus jaculus) and one lagomorph (Ochotona princeps) has prompted us to reinvestigate the issue. Using Bayesian phylogenetics, we found phylogenetic relationships between mammalian species highly congruent with previous results based on nuclear genes. Our results show the existence of four primary lineages of placental mammals: Xenarthra, Afrotheria, Laurasiatheria, and Euarchontoglires. Relationships between and within these lineages strongly suggest that the gene trees may also be congruent with the underlying species phylogeny.

Animals↗

Transcript mapping and genome annotation of ascidian mtDNA using EST data.

Mitochondrial transcripts of two ascidian species were reconstructed through sequence assembly of publicly available ESTs resembling mitochondrial DNA sequences (mt-ESTs). This strategy allowed us to analyze processing and mapping of the mitochondrial transcripts and to investigate the gene organization of a previously uncharacterized mitochondrial genome (mtDNA). This new strategy would greatly facilitate the sequencing and annotation of mtDNAs. In Ciona intestinalis, the assembled mt-ESTs covered 22 mitochondrial genes ( approximately 12,000 bp) and provided the partial sequence of the mtDNA and the prediction of its gene organization. Such sequences were confirmed by amplification and sequencing of the entire Ciona mtDNA. For Halocynthia roretzi, for which the mtDNA sequence was already available, the inferred mt transcripts allowed better definition of gene boundaries (16S rRNA, ND1, ATP6, and tRNA-Ser genes) and the identification of a new gene (an additional Phe-tRNA). In both species, polycistronic and immature transcripts, creation of stop codons by polyadenylation, tRNA signal processing, and rRNA transcript termination signals were identified, thus suggesting that the main features of mitochondrial transcripts are conserved in Chordata.

Animals↗

Computational identification of protein coding potential of conserved sequence tags through cross-species evolutionary analysis.

The identification of conserved sequence tags (CSTs) through comparative genome analysis may reveal important regulatory elements involved in shaping the spatio-temporal expression of genetic information. It is well known that the most significant fraction of CSTs observed in human-mouse comparisons correspond to protein coding exons, due to their strong evolutionary constraints. As we still do not know the complete gene inventory of the human and mouse genomes it is of the utmost importance to establish if detected conserved sequences are genes or not. We propose here a simple algorithm that, based on the observation of the specific evolutionary dynamics of coding sequences, efficiently discriminates between coding and non-coding CSTs. The application of this method may help the validation of predicted genes, the prediction of alternative splicing patterns in known and unknown genes and the definition of a dictionary of non-coding regulatory elements.

Algorithms↗

PatSearch: A program for the detection of patterns and structural motifs in nucleotide sequences.

Regulation of gene expression at transcriptional and post-transcriptional level involves the interaction between short DNA or RNA tracts and the corresponding trans-acting protein factors. Detection of such cis-acting elements in genome-wide screenings may significantly contribute to genome annotation and comparative analysis as well as to target functional characterization experiments. We present here PatSearch, a flexible and fast pattern matcher able to search for specific combinations of oligonucleotide consensus sequences, secondary structure elements and position-weight matrices. It can also allow for mismatches/mispairings below a user fixed threshold. We report three different applications of the program in the search of complex patterns such as those of the iron responsive element hairpin-loop structure, the p53 responsive element and a promoter module containing CAAT-, TATA- and cap-boxes. PatSearch is available on the web at http://bighost.area.ba.cnr.it/BIG/PatSearch/.

Base Sequence↗

Translational control of Scamper expression via a cell-specific internal ribosome entry site.

The mRNA of Scamper, a putative intracellular calcium channel activated by sphingosylphosphocholine, contains a long 5' transcript leader with several upstream AUGs. In this work we have investigated the role this sequence plays in the translational control of Scamper expression. The cytosolic transcription machinery of a T7 RNA polymerase recombinant vaccinia virus was used to avoid artifacts arising from cryptic promoters or mRNA processing. Based on transient transfection experiments of dicistronic and bi-monocistronic plasmids expressing reporter genes, we present evidence that the 5' transcript leader of Scamper contains a functional internal ribosome entry site (IRES). Our data indicate that Scamper translation in Madin-Darby canine kidney cells is driven by a cap-independent mechanism supported by the IRES activity of its mRNA. Finally, the Scamper IRES appears to be the first IRES with specificity for kidney epithelial cells.

5' Untranslated Regions↗