Search PubMed⌕ Search

Biomedical subjects

Luciano Milanesi

Publications and source records attributed to Luciano Milanesi.

14 recordsLinked to original sources

Strategies for comparing gene expression profiles from different microarray platforms: application to a case-control experiment.

Meta-analysis of microarray data is increasingly important, considering both the availability of multiple platforms using disparate technologies and the accumulation in public repositories of data sets from different laboratories. We addressed the issue of comparing gene expression profiles from two microarray platforms by devising a standardized investigative strategy. We tested this procedure by studying MDA-MB-231 cells, which undergo apoptosis on treatment with resveratrol. Gene expression profiles were obtained using high-density, short-oligonucleotide, single-color microarray platforms: GeneChip (Affymetrix) and CodeLink (Amersham). Interplatform analyses were carried out on 8414 common transcripts represented on both platforms, as identified by LocusLink ID, representing 70.8% and 88.6% of annotated GeneChip and CodeLink features, respectively. We identified 105 differentially expressed genes (DEGs) on CodeLink and 42 DEGs on GeneChip. Among them, only 9 DEGs were commonly identified by both platforms. Multiple analyses (BLAST alignment of probes with target sequences, gene ontology, literature mining, and quantitative real-time PCR) permitted us to investigate the factors contributing to the generation of platform-dependent results in single-color microarray experiments. An effective approach to cross-platform comparison involves microarrays of similar technologies, samples prepared by identical methods, and a standardized battery of bioinformatic and statistical analyses.

Breast Neoplasms↗

A novel polymorphism in SEL1L confers susceptibility to Alzheimer's disease.

Alzheimer's disease (AD) is considered to be a conformational disease arising from the accumulation of misfolded and unfolded proteins in the endoplasmic reticulum (ER). SEL1L is a component of the ER stress degradation system, which serves to remove unfolded proteins by retrograde degradation using the ubiquitin-proteosome system. In order to identify genetic variations possibly involved in the disease, we analysed the entire SEL1L gene sequence in Italian sporadic AD patients. Here we report on the identification of a new polymorphism within the SEL1L intron 3 (IVS3-88 A>G), which contains potential binding sites for transcription factors involved in ER-induced stress. Our statistical analysis shows a possible role of the novel polymorphism as independent susceptibility factor of Alzheimer's dementia.

Aged↗

Modelling the interaction of steroid receptors with endocrine disrupting chemicals.

BACKGROUND: The organic polychlorinated compounds like dichlorodiphenyltrichloroethane with its metabolites and polychlorinated biphenyls are a class of highly persistent environmental contaminants. They have been recognized to have detrimental health effects both on wildlife and humans acting as endocrine disrupters due to their ability of mimicking the action of the steroid hormones, and thus interfering with hormone response. There are several experimental evidences that they bind and activate human steroid receptors. However, despite the growing concern about the toxicological activity of endocrine disrupters, molecular data of the interaction of these compounds with biological targets are still lacking. RESULTS: We have used a flexible docking approach to characterize the molecular interaction of seven endocrine disrupting chemicals with estrogen, progesterone and androgen receptors in the ligand-binding domain. All ligands docked in the buried hydrophobic cavity corresponding to the hormone steroid pocket. The interaction was characterized by multiple hydrophobic contacts involving a different number of residues facing the binding pocket, depending on ligands orientation. The EDC ligands did not display a unique binding mode, probably due to their lipophilicity and flexibility, which conferred them a great adaptability into the hydrophobic and large binding pocket of steroid receptors. CONCLUSION: Our results are in agreement with toxicological data on binding and allow to describe a pattern of interactions for a group of ECD to steroid receptors suggesting the requirement of a hydrophobic cavity to accommodate these chlorine carrying compounds. Although the affinity is lower than for hormones, their action can be brought about by a possible synergistic effect.

Amino Acid Sequence↗

ESTree db: a tool for peach functional genomics.

BACKGROUND: The ESTree db http://www.itb.cnr.it/estree/ represents a collection of Prunus persica expressed sequenced tags (ESTs) and is intended as a resource for peach functional genomics. A total of 6,155 successful EST sequences were obtained from four in-house prepared cDNA libraries from Prunus persica mesocarps at different developmental stages. Another 12,475 peach EST sequences were downloaded from public databases and added to the ESTree db. An automated pipeline was prepared to process EST sequences using public software integrated by in-house developed Perl scripts and data were collected in a MySQL database. A php-based web interface was developed to query the database. RESULTS: The ESTree db version as of April 2005 encompasses 18,630 sequences representing eight libraries. Contig assembly was performed with CAP3. Putative single nucleotide polymorphism (SNP) detection was performed with the AutoSNP program and a search engine was implemented to retrieve results. All the sequences and all the contig consensus sequences were annotated both with blastx against the GenBank nr db and with GOblet against the viridiplantae section of the Gene Ontology db. Links to NiceZyme (Expasy) and to the KEGG metabolic pathways were provided. A local BLAST utility is available. A text search utility allows querying and browsing the database. Statistics were provided on Gene Ontology occurrences to assign sequences to Gene Ontology categories. CONCLUSION: The resulting database is a comprehensive resource of data and links related to peach EST sequences. The Sequence Report and Contig Report pages work as the web interface core structures, giving quick access to data related to each sequence/contig.

Chromosome Mapping↗

High performance workflow implementation for protein surface characterization using grid technology.

BACKGROUND: This study concerns the development of a high performance workflow that, using grid technology, correlates different kinds of Bioinformatics data, starting from the base pairs of the nucleotide sequence to the exposed residues of the protein surface. The implementation of this workflow is based on the Italian Grid.it project infrastructure, that is a network of several computational resources and storage facilities distributed at different grid sites. METHODS: Workflows are very common in Bioinformatics because they allow to process large quantities of data by delegating the management of resources to the information streaming. Grid technology optimizes the computational load during the different workflow steps, dividing the more expensive tasks into a set of small jobs. RESULTS: Grid technology allows efficient database management, a crucial problem for obtaining good results in Bioinformatics applications. The proposed workflow is implemented to integrate huge amounts of data and the results themselves must be stored into a relational database, which results as the added value to the global knowledge. CONCLUSION: A web interface has been developed to make this technology accessible to grid users. Once the workflow has started, by means of the simplified interface, it is possible to follow all the different steps throughout the data processing. Eventually, when the workflow has been terminated, the different features of the protein, like the amino acids exposed on the protein surface, can be compared with the data present in the output database.

Automation↗

Systematic analysis of human kinase genes: a large number of genes and alternative splicing events result in functional and structural diversity.

BACKGROUND: Protein kinases are a well defined family of proteins, characterized by the presence of a common kinase catalytic domain and playing a significant role in many important cellular processes, such as proliferation, maintenance of cell shape, apoptosis. In many members of the family, additional non-kinase domains contribute further specialization, resulting in subcellular localization, protein binding and regulation of activity, among others. About 500 genes encode members of the kinase family in the human genome, and although many of them represent well known genes, a larger number of genes code for proteins of more recent identification, or for unknown proteins identified as kinase only after computational studies. RESULTS: A systematic in silico study performed on the human genome, led to the identification of 5 genes, on chromosome 1, 11, 13, 15 and 16 respectively, and 1 pseudogene on chromosome X; some of these genes are reported as kinases from NCBI but are absent in other databases, such as KinBase. Comparative analysis of 483 gene regions and subsequent computational analysis, aimed at identifying unannotated exons, indicates that a large number of kinase may code for alternately spliced forms or be incorrectly annotated. An InterProScan automated analysis was performed to study domain distribution and combination in the various families. At the same time, other structural features were also added to the annotation process, including the putative presence of transmembrane alpha helices, and the cystein propensity to participate into a disulfide bridge. CONCLUSION: The predicted human kinome was extended by identifying both additional genes and potential splice variants, resulting in a varied panorama where functionality may be searched at the gene and protein level. Structural analysis of kinase proteins domains as defined in multiple sources together with transmembrane alpha helices and signal peptide prediction provides hints to function assignment. The results of the human kinome analysis are collected in the KinWeb database, available for browsing and searching over the internet, where all results from the comparative analysis and the gene structure annotation are made available, alongside the domain information. Kinases may be searched by domain combinations and the relative genes may be viewed in a graphic browser at various level of magnification up to gene organization on the full chromosome set.

Algorithms↗

Web services and workflow management for biological resources.

BACKGROUND: The completion of the Human Genome Project has resulted in large quantities of biological data which are proving difficult to manage and integrate effectively. There is a need for a system that is able to automate accesses to remote sites and to "understand" the information that it is managing in order to link data properly. Workflow management systems combined with Web Services are promising Information and Communication Technologies (ICT) tools. Some have already been proposed and are being increasingly applied to the biomedical domain, especially as many biology-related Web Services are now becoming available. Information on biological resources and on genomic sequences mutations are two examples of very specialized datasets that are useful for specific research domains. RESULTS: The architecture of a system that is able to access and execute predefined workflows is presented in this paper. Web Services allowing access to the IARC TP53 Mutation Database and CABRI catalogues of biological resources have been implemented and are available on-line. Example workflows which retrieve data from these Web Services have also been created and are available on-line. CONCLUSION: We present a general architecture and some building blocks for the implementation of a system that is able to remotely execute workflows of biomedical interest and show how this approach can effectively produce useful outputs. The further development and implementation of Web Services allowing access to an exhaustive set of biomedical databases and the creation of effective and useful workflows will improve the automation of in-silico analysis.

Animals↗

A hybrid genetic-neural system for predicting protein secondary structure.

BACKGROUND: Due to the strict relation between protein function and structure, the prediction of protein 3D-structure has become one of the most important tasks in bioinformatics and proteomics. In fact, notwithstanding the increase of experimental data on protein structures available in public databases, the gap between known sequences and known tertiary structures is constantly increasing. The need for automatic methods has brought the development of several prediction and modelling tools, but a general methodology able to solve the problem has not yet been devised, and most methodologies concentrate on the simplified task of predicting secondary structure. RESULTS: In this paper we concentrate on the problem of predicting secondary structures by adopting a technology based on multiple experts. The system performs an overall processing based on two main steps: first, a "sequence-to-structure" prediction is enforced by resorting to a population of hybrid (genetic-neural) experts, and then a "structure-to-structure" prediction is performed by resorting to an artificial neural network. Experiments, performed on sequences taken from well-known protein databases, allowed to reach an accuracy of about 76%, which is comparable to those obtained by state-of-the-art predictors. CONCLUSION: The adoption of a hybrid technique, which encompasses genetic and neural technologies, has demonstrated to be a promising approach in the task of protein secondary structure prediction.

Algorithms↗

Multiple alignment through protein secondary-structure information.

It is well known that protein secondary-structure information can help the process of performing multiple alignment, in particular when the amount of similarity among the involved sequences moves toward the "twilight zone" (less than 30% of pairwise similarity). In this paper, a multiple alignment algorithm is presented, explicitly designed for exploiting any available secondary-structure information. A layered architecture with two interacting levels has been defined for dealing with both primary- and secondary-structure information of target sequences. Secondary structure (either available or predicted by resorting to a technique based on multiple experts) is used to calculate an initial alignment at the secondary level, to be arranged by locally scoped operators devised to refine the alignment at the primary level. Aimed at evaluating the impact of secondary information on the quality of alignments, in particular alignments with a low degree of similarity, the technique has been implemented and assessed on relevant test cases.

Algorithms↗

Representation and modeling of protein surface determinants.

Surface characterization of peptides may provide useful information about functionality and potential interactions with other molecules. A description of a protein site through a surface that models the shape conferred by the exposed residues is an effective tool for the analysis and the modeling of proteins that may highlight similarities and relationships not detectable through comparisons at level of primary, secondary, and tertiary structure. This study concerns the development of a tool that extracts the residues that concur to the shape modeling of the surface of a protein or a portion of it. This task is accomplished without taking into account the order of the amino acids in the primary structure, but only according to the selection of a portion of the protein indicated through geometric parameters or an explicit list of amino acids belonging to the site of interest. Both in the case of an entire protein and in the case of a portion of it, the method provides the mesh that models the surface described by the exposed residues that constitute the external envelope. The developed tool which allows the extraction of the exposed residues, and thus of the potential function determinants, is applied to identify the amino acids that concur to the structural interaction in several protein complexes.

Amino Acid Sequence↗

From context-dependence of mutations to molecular mechanisms of mutagenesis.

Mutation frequencies vary significantly along nucleotide sequences such that mutations often concentrate at certain positions called hotspots. Mutation hotspots in DNA reflect intrinsic properties of the mutation process, such as sequence specificity, that manifests itself at the level of interaction between mutagens, DNA, and the action of the repair and replication machineries. The nucleotide sequence context of mutational hotspots is a fingerprint of interactions between DNA and repair/replication/modification enzymes, and the analysis of hotspot context provides evidence of such interactions. The hotspots might also reflect structural and functional features of the respective DNA sequences and provide information about natural selection. We discuss analysis of 8-oxoguanine-induced mutations in pro- and eukaryotic genes, polymorphic positions in the human mitochondrial DNA and mutations in the HIV-1 retrovirus. Comparative analysis of 8-oxoguanine-induced mutations and spontaneous mutation spectra suggested that a substantial fraction of spontaneous A x T-->C x T mutations is caused by 8-oxoGTP in nucleotide pools. In the case of human mitochondrial DNA, significant differences between molecular mechanisms of mutations in hypervariable segments and coding part of DNA were detected. Analysis of mutations in the HIV-1 retrovirus suggested a complex interplay between molecular mechanisms of mutagenesis and natural selection.

Base Sequence↗

Computational analysis of mutation spectra.

Mutation frequencies vary along a nucleotide sequence, and nucleotide positions with an exceptionally high mutation frequency are called hotspots. Mutation hotspots in DNA often reflect intrinsic properties of the mutation process, such as the specificity with which mutagens interact with nucleic acids and the sequence-specificity of DNA repair/replication enzymes. They might also reflect structural and functional features of target protein or RNA sequences in which they occur. The determinants of mutation frequency and specificity are complex and there are many analytical methods for their study. This paper discusses computational approaches to analysing mutation spectra (distribution of mutations along the target genes) that include many detectable (mutable) positions. The following methods are reviewed: mutation hotspot prediction; pairwise and multiple comparisons of mutation spectra; derivation of a consensus sequence; and analysis of correlation between nucleotide sequence features and mutation spectra. Spectra of spontaneous and induced mutations are used for illustration of the complexities and pitfalls of such analyses. In general, the DNA sequence context of mutation hotspots is a fingerprint of interactions between DNA and DNA repair/replication/modification enzymes, and the analysis of hotspot context provides evidence of such interactions.

Amino Acid Motifs↗

ESTMAP: a system for expressed sequence tags mapping on genomic sequences.

The completion of a number of large genome sequencing projects emphasizes the importance of protein-coding gene predictions. Most of the problems associated with gene prediction are caused by the complex exon-intron structures commonly found in eukaryotic genomes. However, information from homologous sequences can significantly improve the accuracy of the prediction. In particular, expressed sequence tags (ESTs) are very useful for this purpose, since currently existing EST collections are very large. We developed an ESTMAP system, which utilizes homology searches against a database of repetitive elements using the RepeatView program and the EST Division of GenBank using the BLASTN program. ESTMAP extracts "exact" matches with EST sequences (> 95% of homology) from BLASTN output file and predicts introns in DNA comparing ESTs and a query sequence. ESTMAP is implemented as a part of the WebGene system (http://www.cnr.it/webgene).

Base Sequence↗

ASPD (Artificially Selected Proteins/Peptides Database): a database of proteins and peptides evolved in vitro.

ASPD is a new curated database that incorporates data on full-length proteins, protein domains and peptides that were obtained through in vitro directed evolution processes (mainly by means of phage display). At present, the ASPD database contains data on 195 selection experiments, which were described in 112 original papers. For each experiment, the following information is given: (i) description of the target for binding, (ii) description of the protein or peptide which serves as the template for library construction and description of the native protein which binds the target, (iii) links to the major proteomic databases (SWISS-PROT, PDB, PROSITE and ENZYME), (iv) keywords referring to the biological significance of the experiment, (v) aligned sequences of proteins or peptides retrieved through in vitro evolution and relevant native or constructed sequences, (vi) the number of rounds of selection/amplification and (vii) the number of occurrences of clones with each sequence. The literature data include a full reference, a link to the MEDLINE database and the name of the corresponding author with his email address. ASPD has a user-friendly interface which allows for simple queries using the names of proteins and ligands, as well as keywords describing the biological role of the interaction studied, and also for queries based on authors' names. It is also possible to access the database by means of the SRS system, allowing complex queries. There is a BLAST search tool against the ASPD for looking directly for homologous sequences. Research tools of the ASPD allow the analysis of pairwise correlations in the sequences of proteins and peptides selected against one target. The URL for the ASPD database is http://www.sgi.sscc.ru/mgs/gnw/aspd/.

Animals↗