Search PubMed⌕ Search

Biomedical subjects

Kristian Vlahovicek

Publications and source records attributed to Kristian Vlahovicek.

12 recordsLinked to original sources

Comparison of codon usage measures and their applicability in prediction of microbial gene expressivity.

BACKGROUND: There are a number of methods (also called: measures) currently in use that quantify codon usage in genes. These measures are often influenced by other sequence properties, such as length. This can introduce strong methodological bias into measurements; therefore we attempted to develop a method free from such dependencies. One of the common applications of codon usage analyses is to quantitatively predict gene expressivity. RESULTS: We compared the performance of several commonly used measures and a novel method we introduce in this paper--Measure Independent of Length and Composition (MILC). Large, randomly generated sequence sets were used to test for dependence on (i) sequence length, (ii) overall amount of codon bias and (iii) codon bias discrepancy in the sequences. A derivative of the method, named MELP (MILC-based Expression Level Predictor) can be used to quantitatively predict gene expression levels from genomic data. It was compared to other similar predictors by examining their correlation with actual, experimentally obtained mRNA or protein abundances. CONCLUSION: We have established that MILC is a generally applicable measure, being resistant to changes in gene length and overall nucleotide composition, and introducing little noise into measurements. Other methods, however, may also be appropriate in certain applications. Our efforts to quantitatively predict gene expression levels in several prokaryotes and unicellular eukaryotes met with varying levels of success, depending on the experimental dataset and predictor used. Out of all methods, MELP and Rainer Merkl's GCB method had the most consistent behaviour. A 'reference set' containing known ribosomal protein genes appears to be a valid starting point for a codon usage-based expressivity prediction.

Chi-Square Distribution↗

CX, DPX and PRIDE: WWW servers for the analysis and comparison of protein 3D structures.

The WWW servers at http://www.icgeb.org/protein/ are dedicated to the analysis of protein 3D structures submitted by the users as the Protein Data Bank (PDB) files. CX computes an atomic protrusion index that makes it possible to highlight the protruding atoms within a protein 3D structure. DPX calculates a depth index for the buried atoms and makes it possible to analyze the distribution of buried residues. CX and DPX return PDB files containing the calculated indices that can then be visualized using standard programs, such as Swiss-PDBviewer and Rasmol. PRIDE compares 3D structures using a fast algorithm based on the distribution of inter-atomic distances. The options include pairwise as well as multiple comparisons, and fold recognition based on searching the CATH fold database.

Algorithms↗

Efficient recognition of folds in protein 3D structures by the improved PRIDE algorithm.

UNLABELLED: An improved version of the PRIDE (PRobaility of IDEntity) fold prediction algorithm has been developed, based on more solid statistical basis, fast search capabilities and efficient input structure processing. The new algorithm is effective in identifying protein structures at the 'H' level of the CATH hierarchy. AVAILABILITY: The new algorithm is integrated into the PRIDE2 web servers at http://pride.szbk.u-szeged.hu and http://www.icgeb.org/pride. SUPPLEMENTARY INFORMATION: Detailed documentation and performance evaluation is available in the description section of the PRIDE2 web server.

Algorithms↗

The SBASE domain sequence resource, release 12: prediction of protein domain-architecture using support vector machines.

SBASE (http://www.icgeb.trieste.it/sbase) is an online resource designed to facilitate the detection of domain homologies based on sequence database search. The present release of the SBASE A library of protein domain sequences contains 972,397 protein sequence segments annotated by structure, function, ligand-binding or cellular topology, clustered into 8547 domain groups. SBASE B contains 169,916 domain sequences clustered into 2526 less well-characterized groups. Domain prediction is based on an evaluation of database search results in comparison with a 'similarity network' of inter-sequence similarity scores, using support vector machines trained on similarity search results of known domains.

Artificial Intelligence↗

Identification and functional characterization of five novel mutant alleles in 58 Italian patients with Gaucher disease type 1.

Gaucher disease (GD) is the most frequent lysosomal glycolipid storage disorder due to an autosomal recessive deficiency of acid beta-glucosidase characterized by the accumulation of glucocerebroside. In this work we carried out the molecular analysis of the glucocerebrosidase gene (GBA) in 58 unrelated patients with GD type 1. We identified five novel genetic alterations: three missense changes c.187G>A (p.D63N), c.473T>G (p.I158S), c.689T>A (p.V230E), a gene-pseudogene recombinant allele and a non-pseudogene-derived complex allele [c.1379G>A;c.1469A>G] encoding [p.G460D;p.H490R]. All mutant alleles were present as compound heterozygotes in association with c.1226A>G (p.N409S), the most common mutation in GD1. The missense mutant proteins were expressed in vitro in COS-1 cells and analyzed by enzyme activity, protein processing and intracellular localization. Functional studies also included the c.662C>T (p.P221L) mutation recently reported in the Spanish GD population (Montfort et al., 2004). The missense mutant alleles retained an extremely low residual enzyme activity with respect to wild type; the complex allele expressed no activity. Processing of the mutant proteins was unaltered except for c.473T>G which was differently glycosylated due to the exposition of an additional glycosylation site. Immunofluorescence studies showed that protein trafficking into the lysosomes was unaffected in all cases. Finally, the characterization of the novel recombinant allele identified a crossover involving the GBA gene and pseudogene between intron 5 and exon 7.

Adult↗

Molecular analysis of the HEXA gene in Italian patients with infantile and late onset Tay-Sachs disease: detection of fourteen novel alleles.

Tay-Sachs disease (TSD) is a recessively inherited disorder caused by the hexosaminidase A deficiency. We report the molecular characterization performed on 31 Italian patients, 22 with the infantile, acute form of TSD and nine patients with the subacute juvenile form, biochemically classified as B1 Variant. Of the 29 different alleles identified, fourteen were due to 15 novel mutations, two being in-cis on a new complex allele. The new alleles caused four frameshifts, three premature stop codons, three amino acid changes, two amino acid deletions and two splicing alterations. As previously reported, the c.533G>A (p.R178H) mutation was present either in homozygosity or as compound heterozygote, in all the patients with the late onset TSD form (B1 Variant); the allele frequency in this group is discussed by comparison with that found in infantile TSD.

Alleles↗

INCA: synonymous codon usage analysis and clustering by means of self-organizing map.

UNLABELLED: INteractive Codon usage Analysis (INCA) provides an array of features useful in analysis of synonymous codon usage in whole genomes. In addition to computing codon frequencies and several usage indices, such as 'codon bias', effective Nc and CAI, the primary strength of INCA has numerous options for the interactive graphical display of calculated values, thus allowing visual detection of various trends in codon usage. Finally, INCA includes a specific unsupervised neural network algorithm, the self-organizing map, used for gene clustering according to the preferred utilization of codons. AVAILABILITY: INCA is available for the Win32 platform and is free of charge for academic use. For details, visit the web page http://www.bioinfo-hr.org/inca or contact the author directly. SUPPLEMENTARY INFORMATION: Software is accompanied with a user manual and a short tutorial.

Algorithms↗

Highly reactive cysteine residues are part of the substrate binding site of mammalian dipeptidyl peptidases III.

Dipeptidyl peptidase III (DPP III) is a cytosolic zinc-exopeptidase involved in the intracellular protein catabolism of eukaryotes. Although inhibition by thiol reagents is a general feature of DPP III originating from various species, the function of activity important sulfhydryl groups is still inadequately understood. The present study of the reactivity of these groups was undertaken in order to clarify their biological significance. The inactivation kinetics of human and rat DPP III by sulfhydryl reagent p-hydroxy-mercuribenzoate (pHMB) was monitored by determination of the enzyme's residual activity with fluorimetric detection. Inactivation of this human enzyme exhibited pseudo-first-order kinetics, suggesting that all reactive SH-groups have equivalent reactivity, and the second-order rate constant was calculated to be 3523+/-567M(-1)min(-1). Rat DPP III was hyperreactive to pHMB and showed biphasic kinetics indicating two classes of reactive SH-groups. The second-order rate constants of 3540M(-1)s(-1) for slower reacting sulfhydryl, and 21,855M(-1)s(-1) for faster reacting sulfhydryl were obtained from slopes of linear plots of pseudo-first-order constants versus reagent concentration. Peptide substrates protected both mammalian DPPs III from inactivation by pHMB. Physiological concentrations of biological thiols and H(2)O(2) inactivated the rat DPP III. Human enzyme was resistant to H(2)O(2) attack and less affected by reduced glutathione (GSH) than the rat homologue. A significantly lower DPP III level, determined by activity measurement and Western blotting, was found in the cytosols of highly oxygenated rat tissues. These results provide kinetic evidence that cysteine residues are involved in substrate binding of mammalian DPPs III.

Amino Acid Sequence↗

DNA analysis servers: plot.it, bend.it, model.it and IS.

The WWW servers at http://www.icgeb.trieste.it/dna/ are dedicated to the analysis of user-submitted DNA sequences; plot.it creates parametric plots of 45 physicochemical, as well as statistical, parameters; bend.it calculates DNA curvature according to various methods. Both programs provide 1D as well as 2D plots that allow localisation of peculiar segments within the query. The server model.it creates 3D models of canonical or bent DNA starting from sequence data and presents the results in the form of a standard PDB file, directly viewable on the user's PC using any molecule manipulation program. The recently established introns server allows statistical evaluation of introns in various taxonomic groups and the comparison of taxonomic groups in terms of length, base composition, intron type etc. The options include the analysis of splice sites and a probability test for exon-shuffling.

Computer Graphics↗

The DNA secondary structure of the Bacillus subtilis genome.

The entire genomic DNA sequence of the Gram-positive bacterium Bacillus subtilis reported in the SubtiList database has been subjected in this work to a complete bioinformatic analysis of the potential formation of secondary DNA structures such as hairpins and bending. The most significant of these structures have been mapped with respect to their genomic location and compared to those structures already known to have a physiological role, such as the rho-independent transcription terminators. The distribution of these structures along the bacterial chromosome shows two major features: (i). the concentration of the most curved DNA in the intergenic regions rather than within the ORFs, and (ii). a decreasing gradient of large hairpins from the origin towards the terC end of chromosomal DNA replication. Given the increasing biological relevance of secondary DNA structures, these findings should facilitate further studies on the evolution, dynamics and expression of the genetic information stored in bacterial genomes.

Bacillus subtilis↗

The SBASE domain sequence library, release 10: domain architecture prediction.

SBASE (http://www.icgeb.trieste.it/sbase) is an on-line collection of protein domain sequences and related computational tools designed to facilitate detection of domain homologies based on simple database search. The 10th 'jubilee release' of the SBASE library of protein domain sequences contains 1 052 904 protein sequence segments annotated by structure, function, ligand-binding or cellular topology, clustered into over 6000 domain groups. Domain identification and functional prediction are based on a comparison of BLAST search outputs with a knowledge base of biologically significant similarities extracted from known domain groups. The knowledge base is generated automatically for each domain group from the comparison of within-group ('self') and out-of-group ('non-self') similarities. This is a memory-based approach wherein group-specific similarity functions are automatically learned from the database.

Animals↗

The SBASE protein domain library, release 9.0: an online resource for protein domain identification.

SBASE (http://www.icgeb.trieste.it/sbase) is an online resource of protein domain sequences designed to facilitate detection of domain homologies based on a simple database search. The ninth release of the SBASE library of protein domain sequences contains 320 000 annotated structural, functional, ligand-binding and topogenic segments of proteins clustered into over 3481 domain groups and 483 protein families. Domain identification and functional prediction are based on a comparison of BLAST search outputs with a knowledge base of within-group ('self') and out-of-group ('non-self') similarities of the known domain groups. This is a memory-based approach wherein class-specific similarity functions are automatically learned from the database [Stanfill,C. and Waltz,D. (1986) COMMUN: ACM, 29, 1213-1228].

Animals↗