Search PubMed⌕ Search

Biomedical subjects

Sandor Suhai

Publications and source records attributed to Sandor Suhai.

10 recordsLinked to original sources

CAFTAN: a tool for fast mapping, and quality assessment of cDNAs.

BACKGROUND: The German cDNA Consortium has been cloning full length cDNAs and continued with their exploitation in protein localization experiments and cellular assays. However, the efficient use of large cDNA resources requires the development of strategies that are capable of a speedy selection of truly useful cDNAs from biological and experimental noise. To this end we have developed a new high-throughput analysis tool, CAFTAN, which simplifies these efforts and thus fills the gap between large-scale cDNA collections and their systematic annotation and application in functional genomics. RESULTS: CAFTAN is built around the mapping of cDNAs to the genome assembly, and the subsequent analysis of their genomic context. It uses sequence features like the presence and type of PolyA signals, inner and flanking repeats, the GC-content, splice site types, etc. All these features are evaluated in individual tests and classify cDNAs according to their sequence quality and likelihood to have been generated from fully processed mRNAs. Additionally, CAFTAN compares the coordinates of mapped cDNAs with the genomic coordinates of reference sets from public available resources (e.g., VEGA, ENSEMBL). This provides detailed information about overlapping exons and the structural classification of cDNAs with respect to the reference set of splice variants. The evaluation of CAFTAN showed that is able to correctly classify more than 85% of 5950 selected "known protein-coding" VEGA cDNAs as high quality multi- or single-exon. It identified as good 80.6 % of the single exon cDNAs and 85 % of the multiple exon cDNAs. The program is written in Perl and in a modular way, allowing the adoption of this strategy to other tasks like EST-annotation, or to extend it by adding new classification rules and new organism databases as they become available. We think that it is a very useful program for the annotation and research of unfinished genomes. CONCLUSION: CAFTAN is a high-throughput sequence analysis tool, which performs a fast and reliable quality prediction of cDNAs. Several thousands of cDNAs can be analyzed in a short time, giving the curator/scientist a first quick overview about the quality and the already existing annotation of a set of cDNAs. It supports the rejection of low quality cDNAs and helps in the selection of likely novel splice variants, and/or completely novel transcripts for new experiments.

Chromosome Mapping↗

Scrambling of sequence information in collision-induced dissociation of peptides.

Collision-induced dissociation (CID) of protonated YAGFL-NH2 leads to nondirect sequence fragment ions that cannot directly be derived from the primary peptide structure. Experimental and theoretical evidence indicate that primary fragmentation of the intact peptide leads to the linear YAGFLoxa b5 ion with a C-terminal oxazolone ring that is attacked by the N-terminal amino group to induce formation of a cyclic peptide b5 isomer. The latter can undergo various proton transfer reactions and opens up to form something other than the YAGFLoxa linear b5 isomer, leading to scrambling of sequence information in the CID of protonated YAGFL-NH2.

Amino Acid Sequence↗

Binding of gold clusters with DNA base pairs: a density functional study of neutral and anionic GC-Aun and AT-Aun (n = 4, 8) complexes.

Binding of clusters of gold atoms (Au) with the guanine-cytosine (GC) and adenine-thymine (AT) Watson-Crick DNA base pairs was studied using the density functional theory (DFT). Geometries of the neutral GC-Au(n) and AT-Au(n) and the corresponding anionic (GC-Au(n))(-1) and (AT-Au(n))(-1) (n = 4, 8) complexes were fully optimized in different electronic states, that is, singlet and triplet states for the neutral complexes and doublet and quartet states for the anionic complexes, using the B3LYP density functional method. The 6-31+G basis set was used for all atoms except gold. For gold atoms, the Los Alamos effective core potential (ECP) basis set LanL2DZ was employed. Vibrational frequency calculations were performed to ensure that the optimized structures corresponded to potential energy surface minima. The gold clusters around the neutral GC and AT base pairs have a T-shaped structure, which satisfactorily resemble those observed experimentally and in other theoretical studies. However, in anionic GC and AT base pairs, the gold clusters have extended zigzag and T-shaped structures. We found that guanine and adenine have high affinity for Au clusters, with their N3 and N7 sites being preferentially involved in binding with the same. The calculated adiabatic electron affinities (AEAs) of the GC-Au(n)complexes (n = 4, 8) were found to be much larger than those of the isolated base pairs.

Adenine↗

Discovery of two novel, small-molecule inhibitors of DNA methylation.

DNA methyltransferases are promising targets for cancer therapy. In many cancer cells promoters of tumor suppressor genes are hypermethylated, which results in gene inactivation. It has been shown that DNA methyltransferase inhibitors can suppress tumor growth and have significant therapeutic value. However, the established inhibitors are limited in their application due to their substantial cytotoxicity. To discover novel compounds for the inhibition of human DNA methyltransferases, we have screened a set of small molecules available from the NCI database. Using a 3-dimensional model of the human DNA methyltransferase 1 and a modified docking and scoring procedure, we have identified a small list of molecules with high affinities for the active site of the enzyme. The two highest scoring structures were found to inhibit DNA methyltransferase activity in vitro and in vivo. The newly discovered inhibitors validate our screening procedure and also provide a useful basis for further rational drug development.

Catalytic Domain↗

Tuning of retinal twisting in bacteriorhodopsin controls the directionality of the early photocycle steps.

Productive proton pumping by bacteriorhodopsin requires that, after the all-trans to 13-cis photoisomerization of the retinal chromophore, the photocycle proceeds with proton transfer and not with thermal back-isomerization. The question of how the protein controls these events in the active site is addressed here using quantum mechanical/molecular mechanical reaction-path calculations. The results indicate that, while retinal twisting significantly contributes to lowering the barrier for the thermal cis-trans back-isomerization, the rate-limiting barrier for this isomerization is still 5-6 kcal/mol larger than that for the first proton-transfer step. In this way, the retinal twisting is finely tuned so as to store energy to drive the subsequent photocycle while preventing wasteful back-isomerization.

Bacteriorhodopsins↗

Epigenetic reactivation of tumor suppressor genes by a novel small-molecule inhibitor of human DNA methyltransferases.

DNA methylation regulates gene expression in normal and malignant cells. The possibility to reactivate epigenetically silenced genes has generated considerable interest in the development of DNA methyltransferase inhibitors. Here, we provide a detailed characterization of RG108, a novel small molecule that effectively blocked DNA methyltransferases in vitro and did not cause covalent enzyme trapping in human cell lines. Incubation of cells with low micromolar concentrations of the compound resulted in significant demethylation of genomic DNA without any detectable toxicity. Intriguingly, RG108 caused demethylation and reactivation of tumor suppressor genes, but it did not affect the methylation of centromeric satellite sequences. These results establish RG108 as a DNA methyltransferase inhibitor with fundamentally novel characteristics that will be particularly useful for the experimental modulation of epigenetic gene regulation.

Binding Sites↗

High-throughput protein analysis integrating bioinformatics and experimental assays.

The wealth of transcript information that has been made publicly available in recent years requires the development of high-throughput functional genomics and proteomics approaches for its analysis. Such approaches need suitable data integration procedures and a high level of automation in order to gain maximum benefit from the results generated. We have designed an automatic pipeline to analyse annotated open reading frames (ORFs) stemming from full-length cDNAs produced mainly by the German cDNA Consortium. The ORFs are cloned into expression vectors for use in large-scale assays such as the determination of subcellular protein localization or kinase reaction specificity. Additionally, all identified ORFs undergo exhaustive bioinformatic analysis such as similarity searches, protein domain architecture determination and prediction of physicochemical characteristics and secondary structure, using a wide variety of bioinformatic methods in combination with the most up-to-date public databases (e.g. PRINTS, BLOCKS, INTERPRO, PROSITE SWISSPROT). Data from experimental results and from the bioinformatic analysis are integrated and stored in a relational database (MS SQL-Server), which makes it possible for researchers to find answers to biological questions easily, thereby speeding up the selection of targets for further analysis. The designed pipeline constitutes a new automatic approach to obtaining and administrating relevant biological data from high-throughput investigations of cDNAs in order to systematically identify and characterize novel genes, as well as to comprehensively describe the function of the encoded proteins.

Automation↗

cDNA2Genome: a tool for mapping and annotating cDNAs.

BACKGROUND: In the last years several high-throughput cDNA sequencing projects have been funded worldwide with the aim of identifying and characterizing the structure of complete novel human transcripts. However some of these cDNAs are error prone due to frameshifts and stop codon errors caused by low sequence quality, or to cloning of truncated inserts, among other reasons. Therefore, accurate CDS prediction from these sequences first require the identification of potentially problematic cDNAs in order to speed up the posterior annotation process. RESULTS: cDNA2Genome is an application for the automatic high-throughput mapping and characterization of cDNAs. It utilizes current annotation data and the most up to date databases, especially in the case of ESTs and mRNAs in conjunction with a vast number of approaches to gene prediction in order to perform a comprehensive assessment of the cDNA exon-intron structure. The final result of cDNA2Genome is an XML file containing all relevant information obtained in the process. This XML output can easily be used for further analysis such us program pipelines, or the integration of results into databases. The web interface to cDNA2Genome also presents this data in HTML, where the annotation is additionally shown in a graphical form. cDNA2Genome has been implemented under the W3H task framework which allows the combination of bioinformatics tools in tailor-made analysis task flows as well as the sequential or parallel computation of many sequences for large-scale analysis. CONCLUSIONS: cDNA2Genome represents a new versatile and easily extensible approach to the automated mapping and annotation of human cDNAs. The underlying approach allows sequential or parallel computation of sequences for high-throughput analysis of cDNAs.

Chromosome Mapping↗

Establishment and functional validation of a structural homology model for human DNA methyltransferase 1.

Changes in DNA methylation patterns play an important role in tumorigenesis. The DNA methyltransferase 1 (DNMT1) protein represents a major DNA methyltransferase activity in human cells and is therefore a prominent target for experimental cancer therapies. However, there are only few available inhibitors and their high toxicity and low specificity have so far precluded their broad use in chemotherapy. Based on the strong conservation of catalytic DNA methyltransferase domains we have used a homology modeling approach to determine the three-dimensional structure of the DNMT1 catalytic domain. Our results suggest an overall structural conservation with other DNA methyltransferases but also indicate local conformational differences. To prove the validity of our model we used it as a template to design a novel derivative of the known DNA methyltransferase inhibitor 5-azacytidine. The resulting compound (N4-fluoroacetyl-5-azacytidine) functioned as an efficient inhibitor of DNA methylation in human tumor cell lines and also provides novel opportunities for pharmacological applications.

Amino Acid Sequence↗

PATH: a task for the inference of phylogenies.

UNLABELLED: Phylogenetic Analysis Task in Husar (PATH) is a task for the inference of phylogenies. It executes three phylogenetic methods and automatically chooses the evolutionary model for each set of data. The output of the tasks shows the consensus trees together with full results obtained from all executed methods. AVAILABILITY: PATH is available at the German EMBnet node after registration via www at http://genome.dkfz-heidelberg.de

Databases, Nucleic Acid↗