Search PubMed⌕ Search

Biomedical subjects

P Carpena

Publications and source records attributed to P Carpena.

4 recordsLinked to original sources

Isochore chromosome maps of eukaryotic genomes.

Analytical DNA ultracentrifugation revealed that eukaryotic genomes are mosaics of isochores: long DNA segments (>>300 kb on average) relatively homogeneous in G+C. Important genome features are dependent on this isochore structure, e.g. genes are found predominantly in the GC-richest isochore classes. However, no reliable method is available to rigorously partition the genome sequence into relatively homogeneous regions of different composition, thereby revealing the isochore structure of chromosomes at the sequence level. Homogeneous regions are currently ascertained by plain statistics on moving windows of arbitrary length, or simply by eye on G+C plots. On the contrary, the entropic segmentation method is able to divide a DNA sequence into relatively homogeneous, statistically significant domains. An early version of this algorithm only produced domains having an average length far below the typical isochore size. Here we show that an improved segmentation method, specifically intended to determine the most statistically significant partition of the sequence at each scale, is able to identify the boundaries between long homogeneous genome regions displaying the typical features of isochores. The algorithm precisely locates classes II and III of the human major histocompatibility complex region, two well-characterized isochores at the sequence level, the boundary between them being the first isochore boundary experimentally characterized at the sequence level. The analysis is then extended to a collection of human large contigs. The relatively homogeneous regions we find show many of the features (G+C range, relative proportion of isochore classes, size distribution, and relationship with gene density) of the isochores identified through DNA centrifugation. Isochore chromosome maps, with many potential applications in genomics, are then drawn for all the completely sequenced eukaryotic genomes available.

Animals↗

Effect of trends on detrended fluctuation analysis.

Detrended fluctuation analysis (DFA) is a scaling analysis method used to estimate long-range power-law correlation exponents in noisy signals. Many noisy signals in real systems display trends, so that the scaling results obtained from the DFA method become difficult to analyze. We systematically study the effects of three types of trends--linear, periodic, and power-law trends, and offer examples where these trends are likely to occur in real data. We compare the difference between the scaling results for artificially generated correlated noise and correlated noise with a trend, and study how trends lead to the appearance of crossovers in the scaling behavior. We find that crossovers result from the competition between the scaling of the noise and the "apparent" scaling of the trend. We study how the characteristics of these crossovers depend on (i) the slope of the linear trend; (ii) the amplitude and period of the periodic trend; (iii) the amplitude and power of the power-law trend, and (iv) the length as well as the correlation properties of the noise. Surprisingly, we find that the crossovers in the scaling of noisy signals with trends also follow scaling laws--i.e., long-range power-law dependence of the position of the crossover on the parameters of the trends. We show that the DFA result of noise with a trend can be exactly determined by the superposition of the separate results of the DFA on the noise and on the trend, assuming that the noise and the trend are not correlated. If this superposition rule is not followed, this is an indication that the noise and the superposed trend are not independent, so that removing the trend could lead to changes in the correlation properties of the noise. In addition, we show how to use DFA appropriately to minimize the effects of trends, how to recognize if a crossover indicates indeed a transition from one type to a different type of underlying correlation, or if the crossover is due to a trend without any transition in the dynamical properties of the noise.

Analysis of Variance↗

Finding borders between coding and noncoding DNA regions by an entropic segmentation method.

We present a new computational approach to finding borders between coding and noncoding DNA. This approach has two features: (i) DNA sequences are described by a 12-letter alphabet that captures the differential base composition at each codon position, and (ii) the search for the borders is carried out by means of an entropic segmentation method which uses only the general statistical properties of coding DNA. We find that this method is highly accurate in finding borders between coding and noncoding regions and requires no "prior training" on known data sets. Our results appear to be more accurate than those obtained with moving windows in the discrimination of coding from noncoding DNA.

DNA↗