Search PubMed⌕ Search

Biomedical subjects

Thomas M A Fink

Publications and source records attributed to Thomas M A Fink.

5 recordsLinked to original sources

Identifying genes from up-down properties of microarray expression series.

MOTIVATION: We consider any collection of microarrays that can be ordered to form a progression; for example, as a function of time, severity of disease or dose of a stimulant. By plotting the expression level of each gene as a function of time, or severity, or dose, we form an expression series, or curve, for each gene. While most of these curves will exhibit random fluctuations, some will contain a pattern, and these are the genes that are most likely associated with the quantity used to order them. RESULTS: We introduce a method of identifying the pattern and hence genes in microarray expression curves without knowing what kind of pattern to look for. Key to our approach is the sequence of ups and downs formed by pairs of consecutive data points in each curve. As a benchmark, we blindly identified genes from yeast cell cycles without selecting for periodic or any other anticipated behaviour. CONTACT: tmf20@cam.ac.uk SUPPLEMENTARY INFORMATION: The complete versions of Table 2 and Figure 4, as well as other material, can be found at http://www.lps.ens.fr/~willbran/up-down/ or http://www.tcm.phy.cam.ac.uk/~tmf20/up-down/

Algorithms↗

Characterization of the probabilistic traveling salesman problem.

We show that stochastic annealing can be successfully applied to gain new results on the probabilistic traveling salesman problem. The probabilistic "traveling salesman" must decide on an a priori order in which to visit n cities (randomly distributed over a unit square) before learning that some cities can be omitted. We find the optimized average length of the pruned tour follows E(L(pruned))=sqrt[np](0.872-0.105p)f(np), where p is the probability of a city needing to be visited, and f(np)-->1 as np--> infinity. The average length of the a priori tour (before omitting any cities) is found to follow E(L(a priori))=sqrt[n/p]beta(p), where beta(p)=1/[1.25-0.82 ln(p)] is measured for 0.05< or =p< or =0.6. Scaling arguments and indirect measurements suggest that beta(p) tends towards a constant for p<0.03. Our stochastic annealing algorithm is based on limited sampling of the pruned tour lengths, exploiting the sampling error to provide the analog of thermal fluctuations in simulated (thermal) annealing. The method has general application to the optimization of functions whose cost to evaluate rises with the precision required.

Journal Article↗

Stochastic annealing.

We show how to simulate a system in thermal equilibrium when the energy cannot be evaluated exactly: the error distribution needs to be symmetric, but it does not need to be known. We also solve the Ceperley-Dewing version of this problem, where the error distribution is taken to be fully known. These underlying ideas give an effective optimization strategy for problems where the evaluation of each design can be sampled only statistically, including an application to protein folding.

Models, Statistical↗

Protein design depends on the size of the amino acid alphabet.

We consider the design of proteins to be simultaneously thermodynamically stable in multiple independent and correlated conformations. We first show that a protein can be trained to fold to multiple independent conformations and calculate its capacity. The number of configurations that it can remember is proportional to the logarithm of the number of amino acid species A, independent of chain length. Next we investigate the recognition of correlated conformations, which we apply to funnel design around a single configuration. The maximum basin of attraction, as parametrized in our model, also depends on the number of amino acid species as ln A. We argue that the extent to which the protein energy landscape can be manipulated is fixed, effecting a trade off between well breadth, well depth, and well number. This emerging picture motivates a clearer understanding of the scope and limits of protein and heteropolymer function.

Amino Acids↗

Sequence determination from overlapping fragments: a simple model of whole-genome shotgun sequencing.

Assembling fragments randomly sampled from along a sequence is the basis of whole-genome shotgun sequencing, a technique used to map the DNA of the human and other genomes. We calculate the probability that a random sequence can be recovered from a collection of overlapping fragments. We provide an exact solution for an infinite alphabet and in the case of constant overlaps. For the general problem we apply two assembly strategies and give the probability that the assembly puzzle can be solved in the limit of infinitely many fragments.

Animals↗