Search PubMed⌕ Search

Biomedical subjects

I Grosse

Publications and source records attributed to I Grosse.

8 recordsLinked to original sources

Statistical analysis of the DNA sequence of human chromosome 22.

We study statistical patterns in the DNA sequence of human chromosome 22, the first completely sequenced human chromosome. We find that (i). the 33.4 x 10(6) nucleotide long human chromosome exhibits long-range power-law correlations over more than four orders of magnitude, (ii). the entropies H(n) of the frequency distribution of oligonucleotides of length n (n-mers) grow sublinearly with increasing n, indicating the presence of higher-order correlations for all of the studied lengths 1<or=n<or=10, and (iii). the generalized entropies H(n)(q) of n-mers decrease monotonically with increasing q and the decay of H(n)(q) with q becomes steeper with increasing n<or=10, indicating that the frequency distribution of oligonucleotides becomes increasingly nonuniform as the length n increases. We investigate to what degree known biological features may explain the observed statistical patterns. We find that (iv). the presence of interspersed repeats may cause the sublinear increase of H(n) with n, and that (v). the presence of monomeric tandem repeats as well as the suppression of CG dinucleotides may cause the observed decay of H(n)(q) with q.

Algorithms↗

Computational identification of promoters and first exons in the human genome.

The identification of promoters and first exons has been one of the most difficult problems in gene-finding. We present a set of discriminant functions that can recognize structural and compositional features such as CpG islands, promoter regions and first splice-donor sites. We explain the implementation of the discriminant functions into a decision tree that constitutes a new program called FirstEF. By using different models to predict CpG-related and non-CpG-related first exons, we showed by cross-validation that the program could predict 86% of the first exons with 17% false positives. We also demonstrated the prediction accuracy of FirstEF at the genome level by applying it to the finished sequences of human chromosomes 21 and 22 as well as by comparing the predictions with the locations of the experimentally verified first exons. Finally, we present the analysis of the predicted first exons for all of the 24 chromosomes of the human genome.

Chromosomes, Human, Pair 21↗

Optimization of coding potentials using positional dependence of nucleotide frequencies.

We study the coding potential of human DNA sequences, using the positional asymmetry function (D(p)) and the positional information function (I(q)). Both D(p)and I(q)are based on the positional dependence of single nucleotide frequencies. We investigate the accuracy of D(p)and I(q)in distinguishing coding and non-coding DNA as a function of the parameters p and q, respectively, and explore at which parameters p(opt)and q(opt)both D(p)and I(q)distinguish coding and non-coding DNA most accurately. We compare our findings with classically used parameter values and find that optimized coding potentials yield comparable accuracies as classical frame-independent coding potentials trained on prior data. We find that p(opt)and q(opt)vary only slightly with the sequence length.

Codon↗

Finding borders between coding and noncoding DNA regions by an entropic segmentation method.

We present a new computational approach to finding borders between coding and noncoding DNA. This approach has two features: (i) DNA sequences are described by a 12-letter alphabet that captures the differential base composition at each codon position, and (ii) the search for the borders is carried out by means of an entropic segmentation method which uses only the general statistical properties of coding DNA. We find that this method is highly accurate in finding borders between coding and noncoding regions and requires no "prior training" on known data sets. Our results appear to be more accurate than those obtained with moving windows in the discrimination of coding from noncoding DNA.

DNA↗

Are noncoding sequences of Rickettsia prowazekii remnants of "neutralized" genes?

It has been hypothesized that a large fraction of 24% noncoding DNA in R. prowazekii consists of degraded genes. This hypothesis has been based on the relatively high G+C content of noncoding DNA. However, a comparison with other genomes also having a low overall G+C content shows that this argument would also apply to other bacteria. To test this hypothesis, we study the coding potential in sets of genes, pseudogenes, and intergenic regions. We find that the correlation function and the chi(2)-measure are clearly indicative of the coding function of genes and pseudogenes. However, both coding potentials make almost no indication of a preexisting reading frame in the remaining 23% of noncoding DNA. We simulate the degradation of genes due to single-nucleotide substitutions and insertions/deletions and quantify the number of mutations required to remove indications of the reading frame. We discuss a reduced selection pressure as another possible origin of this comparatively large fraction of noncoding sequences.

DNA, Intergenic↗

Species independence of mutual information in coding and noncoding DNA.

We explore if there exist universal statistical patterns that are different in coding and noncoding DNA and can be found in all living organisms, regardless of their phylogenetic origin. We find that (i) the mutual information function [symbol: see text] has a significantly different functional form in coding and noncoding DNA. We further find that (ii) the probability distributions of the average mutual information [symbol: see text] are significantly different in coding and noncoding DNA, while (iii) they are almost the same for organisms of all taxonomic classes. Surprisingly, we find that [symbol: see text] is capable of predicting coding regions as accurately as organism-specific coding measures.

DNA↗

Scale invariance and lack of self-averaging in fragmentation

We derive exact statistical properties of a recursive fragmentation process. We show that introducing a fragmentation probability 0<p<1 leads to a purely algebraic size distribution, P(x) approximately x(-2p), in one dimension. In d dimensions, the volume distribution diverges algebraically in the small fragment limit, P(V) approximately V-gamma, with gamma=2p(1/d). Hence, the entire range of exponents allowed by mass conservation is realized. We demonstrate that this fragmentation process is non-self-averaging as the moments Y(alpha)= summation operator(i)x(alpha)(i) exhibit significant sample to sample fluctuations.

Journal Article↗

Average mutual information of coding and noncoding DNA.

One basic problem in the analysis of DNA sequences is the recognition of protein-coding genes. Computer algorithms to facilitate gene identification have become important as genome sequencing projects have turned from mapping to large-scale sequencing, resulting in an exponentially growing number of sequenced nucleotides that await their annotation. Many statistical patterns have been discovered that are different in coding and noncoding DNA, but most of them vary from species to species, and hence require prior training on organism-specific data sets. Here, we investigate if there exist species-independent statistical patterns that are different in coding and noncoding DNA. We introduce an information-theoretic quantity, the average mutual information (AMI), and we find that the probability distribution functions of the AMI are significantly different in coding and noncoding DNA, while they are almost identical for different species. This finding suggests that the AMI might be useful for the recognition of protein-coding regions in genomes for which training sets do not exist.

Algorithms↗