Search PubMed⌕ Search

Biomedical subjects

A O Schmitt

Publications and source records attributed to A O Schmitt.

6 recordsLinked to original sources

An algorithm for clustering cDNA fingerprints.

Clustering large data sets is a central challenge in gene expression analysis. The hybridization of synthetic oligonucleotides to arrayed cDNAs yields a fingerprint for each cDNA clone. Cluster analysis of these fingerprints can identify clones corresponding to the same gene. We have developed a novel algorithm for cluster analysis that is based on graph theoretic techniques. Unlike other methods, it does not assume that the clusters are hierarchically structured and does not require prior knowledge on the number of clusters. In tests with simulated libraries the algorithm outperformed the Greedy method and demonstrated high speed and robustness to high error rate. Good solution quality was also obtained in a blind test on real cDNA fingerprints.

Algorithms↗

Differential gene expression by endothelial cells in distinct angiogenic states.

Angiogenesis is a complex process that can be regarded as a series of sequential events comprising a variety of tissue cells. The major problem when studying angiogenesis in vitro is the lack of a model system mimicking the various aspects of the process in vivo. In this study we have used two in vitro models, each representing different and distinct aspects of angiogenesis. Differentially expressed genes in the two culture forms were identified using the suppression subtractive hybridization technique to prepare subtracted cDNA libraries. This was followed by a differential hybridization screen to pick up overexpressed clones. Using comparative multiplex RT-PCR we confirmed the differential expression and showed differences up to 14-fold. We identified a broad range of genes already known to play an important role during angiogenesis like Flt1 or TIE2. Furthermore several known genes are put into the context of endothelial cell differentiation, which up to now have not been described as being relevant to angiogenesis, like NrCAM, Claudin14, BMP-6, PEA-15 and PINCH. With ADAMTS4 and hADAMTS1/METH-1 we further extended the set of matrix metalloproteases expressed and regulated by endothelial cells.

Base Sequence↗

Information theoretical probe selection for hybridisation experiments.

MOTIVATION: The choice of probes is an important feature of hybridisation experiments. In this paper we present an algorithm that optimises probes with respect to a training set of sequences based on Shannon entropy as a quality criterion. The practical motivation for our algorithm is oligonucleotide fingerprinting, a method for the simultaneous identification of sequences (cDNA or genomic DNA) by their hybridisation tags according to a set of short probes such as octamers, although the algorithm is of course not restricted to that application. RESULTS: We can show that our method is superior to the selection of probes according to their frequencies, which is a widely used strategy, and to randomly chosen probe sets. The quality of probe sets is assessed by a simulation pipeline that entails the set of probes as a simulation parameter. The performance of probe sets trained on sequences from different organisms shows additionally that probes should be chosen with regard to the organism under analysis. Case studies are presented on how constraints (G+C-content, complexity of the individual probes) influence the selection process. AVAILABILITY: A description of the oligonucleotide fingerprinting pipeline is published on our web-page http://www.molgen.mpg.de/ approximately ag_onf/met.htm. An executable of the algorithm and probe lists designed for human and rodents can be downloaded from the ftp-site ftp://ftp.molgen.mpg.de/pub/mpimg/probe_design/.

Algorithms↗

Exhaustive mining of EST libraries for genes differentially expressed in normal and tumour tissues.

A four-step procedure for the efficient and systematic mining of whole EST libraries for differentially expressed genes is presented. After eliminating redundant entries from the EST library under investigation (step 1), contigs of maximal length are built upon each remaining EST using about 4 000 000 public and proprietary ESTs (step 2). These putative genes are compared against a database comprising ESTs from 16 different tissues (both normal and tumour affected) to determine whether or not they are differentially expressed (step 3; electronic northern). Fisher's exact test is used to assess the significance of differential expression. In step 4, an attempt is made to characterise the contigs obtained in the assembly through database comparison. A case study of the CGAP library NCI_CGAP_Br1.1, a library made from three (well, moderately, and poorly differentiated) invasive ductal breast tumours (2126 ESTs in total) was carried out. Of the maximal contigs, 139 were found to be significantly (alpha = 0.05) over-expressed in breast tumour tissue, while 13 appeared to be down-regulated.

Animals↗

Estimating the entropy of DNA sequences.

The Shannon entropy is a standard measure for the order state of symbol sequences, such as, for example, DNA sequences. In order to incorporate correlations between symbols, the entropy of n-mers (consecutive strands of n symbols) has to be determined. Here, an assay is presented to estimate such higher order entropies (block entropies) for DNA sequences when the actual number of observations is small compared with the number of possible outcomes. The n-mer probability distribution underlying the dynamical process is reconstructed using elementary statistical principles: The theorem of asymptotic equi-distribution and the Maximum Entropy Principle. Constraints are set to force the constructed distributions to adopt features which are characteristic for the real probability distribution. From the many solutions compatible with these constraints the one with the highest entropy is the most likely one according to the Maximum Entropy Principle. An algorithm performing this procedure is expounded. It is tested by applying it to various DNA model sequences whose exact entropies are known. Finally, results for a real DNA sequence, the complete genome of the Epstein Barr virus, are presented and compared with those of other information carriers (texts, computer source code, music). It seems as if DNA sequences possess much more freedom in the combination of the symbols of their alphabet than written language or computer source codes.

Algorithms↗

The modular structure of informational sequences.

It is shown that DNA sequences can be decomposed into smaller units much the same as texts can be decomposed into syllables, words, or groups of words. Those smaller units (modules) are extracted from DNA sequences according to statistical criteria. Tests with sequences of known modular structure (two novels and a FORTRAN source code) were performed. The rate to which DNA sequences can be decomposed into modules (modularity) turns out to be a very sensitive measure to distinguish DNA sequences from random sequences.

Algorithms↗