Search PubMed⌕ Search

Biomedical subjects

Lorenz Wernisch

Publications and source records attributed to Lorenz Wernisch.

14 recordsLinked to original sources

Microarray analysis after RNA amplification can detect pronounced differences in gene expression using limma.

BACKGROUND: RNA amplification is necessary for profiling gene expression from small tissue samples. Previous studies have shown that the T7 based amplification techniques are reproducible but may distort the true abundance of targets. However, the consequences of such distortions on the ability to detect biological variation in expression have not been explored sufficiently to define the true extent of usability and limitations of such amplification techniques. RESULTS: We show that expression ratios are occasionally distorted by amplification using the Affymetrix small sample protocol version 2 due to a disproportional shift in intensity across biological samples. This occurs when a shift in one sample cannot be reflected in the other sample because the intensity would lie outside the dynamic range of the scanner. Interestingly, such distortions most commonly result in smaller ratios with the consequence of reducing the statistical significance of the ratios. This becomes more critical for less pronounced ratios where the evidence for differential expression is not strong. Indeed, statistical analysis by limma suggests that up to 87% of the genes with the largest and therefore most significant ratios (p < 10e(-20)) in the unamplified group have a p-value below 10e(-20) in the amplified group. On the other hand, only 69% of the more moderate ratios (10e(-20) < p < 10e(-10)) in the unamplified group have a p-value below 10e(-10) in the amplified group. Our analysis also suggests that, overall, limma shows better overlap of genes found to be significant in the amplified and unamplified groups than the Z-scores statistics. CONCLUSION: We conclude that microarray analysis of amplified samples performs best at detecting differences in gene expression, when these are large and when limma statistics are used.

Animals↗

Archaeology and evolution of transfer RNA genes in the Escherichia coli genome.

Transfer RNA genes tend to be presented in multiple copies in the genomes of most organisms, from bacteria to eukaryotes. The evolution and genomic structure of tRNA genes has been a somewhat neglected area of molecular evolution. Escherichia coli, the first phylogenetic species for which more than two different strains have been sequenced, provides an invaluable framework to study the evolution of tRNA genes. In this work, a detailed analysis of the tRNA structure of the genomes of Escherichia coli strains K12, CFT073, and O157:H7, Shigella flexneri 2a 301, and Salmonella typhimurium LT2 was carried out. A phylogenetic analysis of these organisms was completed, and an archaeological map depicting the main events in the evolution of tRNA genes was drawn. It is shown that duplications, deletions, and horizontal gene transfers are the main factors driving tRNA evolution in these genomes. On average, 0.64 tRNA insertions/duplications occur every million years (Myr) per genome per lineage, while deletions occur at the slower rate of 0.30 per million years per genome per lineage. This work provides a first genomic glance at the problem of tRNA evolution as a repetitive process, and the relationship of this mechanism to genome evolution and codon usage is discussed.

Codon↗

A Hidden Markov model web application for analysing bacterial genomotyping DNA microarray experiments.

Whole genome DNA microarray genomotyping experiments compare the gene content of different species or strains of bacteria. A statistical approach to analysing the results of these experiments was developed, based on a Hidden Markov model (HMM), which takes adjacency of genes along the genome into account when calling genes present or absent. The model was implemented in the statistical language R and applied to three datasets. The method is numerically stable with good convergence properties. Error rates are reduced compared with approaches that ignore spatial information. Moreover, the HMM circumvents a problem encountered in a conventional analysis: determining the cut-off value to use to classify a gene as absent. An Apache Struts web interface for the R script was created for the benefit of users unfamiliar with R. The application may be found at http://hmmgd.cryst.bbk.ac.uk/hmmgd. The source code illustrating how to run R scripts from an Apache Struts-based web application is available from the corresponding author on request. The application is also available for local installation if required.

Algorithms↗

Applying GIFT, a Gene Interactions Finder in Text, to fly literature.

UNLABELLED: A number of freely available text mining tools have been put together to extract highly reliable Drosophila gene interaction data from text. The system has been tested with The Interactive Fly, showing low recall (27-34%), but very high precision (93-97%). AVAILABILITY: The extracted data and a web interface for submission of texts to GIFT analysis are available at http://gift.cryst.bbk.ac.uk/gift CONTACT: n.domedel_puig@cryst.bbk.ac.uk SUPPLEMENTARY INFORMATION: Additional documentation, such as the dictionaries and the reference sets, are available at the GIFT website.

Artificial Intelligence↗

A universally applicable method of operon map prediction on minimally annotated genomes using conserved genomic context.

An important step in understanding the regulation of a prokaryotic genome is the generation of its transcription unit map. The current strongest operon predictor depends on the distributions of intergenic distances (IGD) separating adjacent genes within and between operons. Unfortunately, experimental data on these distance distributions are limited to Escherichia coli and Bacillus subtilis. We suggest a new graph algorithmic approach based on comparative genomics to identify clusters of conserved genes independent of IGD and conservation of gene order. As a consequence, distance distributions of operon pairs for any arbitrary prokaryotic genome can be inferred. For E.coli, the algorithm predicts 854 conserved adjacent pairs with a precision of 85%. The IGD distribution for these pairs is virtually identical to the E.coli operon pair distribution. Statistical analysis of the predicted pair IGD distribution allows estimation of a genome-specific operon IGD cut-off, obviating the requirement for a training set in IGD-based operon prediction. We apply the method to a representative set of eight genomes, and show that these genome-specific IGD distributions differ considerably from each other and from the distribution in E.coli.

Algorithms↗

Solving the riddle of codon usage preferences: a test for translational selection.

Translational selection is responsible for the unequal usage of synonymous codons in protein coding genes in a wide variety of organisms. It is one of the most subtle and pervasive forces of molecular evolution, yet, establishing the underlying causes for its idiosyncratic behaviour across living kingdoms has proven elusive to researchers over the past 20 years. In this study, a statistical model for measuring translational selection in any given genome is developed, and the test is applied to 126 fully sequenced genomes, ranging from archaea to eukaryotes. It is shown that tRNA gene redundancy and genome size are interacting forces that ultimately determine the action of translational selection, and that an optimal genome size exists for which this kind of selection is maximal. Accordingly, genome size also presents upper and lower boundaries beyond which selection on codon usage is not possible. We propose a model where the coevolution of genome size and tRNA genes explains the observed patterns in translational selection in all living organisms. This model finally unifies our understanding of codon usage across prokaryotes and eukaryotes. Helicobacter pylori, Saccharomyces cerevisiae and Homo sapiens are codon usage paradigms that can be better understood under the proposed model.

Animals↗

Predicting metal-binding site residues in low-resolution structural models.

The accurate prediction of the biochemical function of a protein is becoming increasingly important, given the unprecedented growth of both structural and sequence databanks. Consequently, computational methods are required to analyse such data in an automated manner to ensure genomes are annotated accurately. Protein structure prediction methods, for example, are capable of generating approximate structural models on a genome-wide scale. However, the detection of functionally important regions in such crude models, as well as structural genomics targets, remains an extremely important problem. The method described in the current study, MetSite, represents a fully automatic approach for the detection of metal-binding residue clusters applicable to protein models of moderate quality. The method involves using sequence profile information in combination with approximate structural data. Several neural network classifiers are shown to be able to distinguish metal sites from non-sites with a mean accuracy of 94.5%. The method was demonstrated to identify metal-binding sites correctly in LiveBench targets where no obvious metal-binding sequence motifs were detectable using InterPro. Accurate detection of metal sites was shown to be feasible for low-resolution predicted structures generated using mGenTHREADER where no side-chain information was available. High-scoring predictions were observed for a recently solved hypothetical protein from Haemophilus influenzae, indicating a putative metal-binding site.

Binding Sites↗

Reconstruction of gene networks using Bayesian learning and manipulation experiments.

MOTIVATION: The analysis of high-throughput experimental data, for example from microarray experiments, is currently seen as a promising way of finding regulatory relationships between genes. Bayesian networks have been suggested for learning gene regulatory networks from observational data. Not all causal relationships can be inferred from correlation data alone. Often several equivalent but different directed graphs explain the data equally well. Intervention experiments where genes are manipulated can help to narrow down the range of possible networks. RESULTS: We describe an active learning algorithm that suggests an optimized sequence of intervention experiments. Simulation experiments show that our selection scheme is better than an unguided choice of interventions in learning the correct network and compares favorably in running time and results with methods based on value of information calculations.

Algorithms↗

The influence of reduced oxygen availability on pathogenicity and gene expression in Mycobacterium tuberculosis.

We investigated how Mycobacterium tuberculosis responded to a reduced oxygen tension in terms of its pathogenicity and gene expression by growing cells under either aerobic or low-oxygen conditions in chemostat culture. The chemostat enabled us to control and vary the oxygen tension independently of other environmental parameters, so that true cause-and-effect relationships of reduced oxygen availability could be established. Cells grown under low oxygen were more pathogenic for guinea pigs than those grown aerobically. The effect of reduced oxygen on global gene expression was determined using DNA microarray. Spearman rank correlation confirmed that microarray expression profiles were highly reproducible between repeat cultures. Using microarray analysis we have identified genes that respond to a low-oxygen environment without the influence of other parameters such as nutrient depletion. Some of these genes appear to be involved in the biosynthesis of cell wall precursors and their induction may have contributed to increased infectivity in the guinea pig. This study has shown that a combination of chemostat culture and microarray presents a biologically robust and statistically reliable experimental approach for studying the effect of relevant and specific environmental stimuli on mycobacterial virulence and gene expression.

Anaerobiosis↗

Unexpected correlations between gene expression and codon usage bias from microarray data for the whole Escherichia coli K-12 genome.

Escherichia coli has long been regarded as a model organism in the study of codon usage bias (CUB). However, most studies in this organism regarding this topic have been computational or, when experimental, restricted to small datasets; particularly poor attention has been given to genes with low CUB. In this work, correspondence analysis on codon usage is used to classify E.coli genes into three groups, and the relationship between them and expression levels from microarray experiments is studied. These groups are: group 1, highly biased genes; group 2, moderately biased genes; and group 3, AT-rich genes with low CUB. It is shown that, surprisingly, there is a negative correlation between codon bias and expression levels for group 3 genes, i.e. genes with extremely low codon adaptation index (CAI) values are highly expressed, while group 2 show the lowest average expression levels and group 1 show the usual expected positive correlation between CAI and expression. This trend is maintained over all functional gene groups, seeming to contradict the E.coli-yeast paradigm on CUB. It is argued that these findings are still compatible with the mutation-selection balance hypothesis of codon usage and that E.coli genes form a dynamic system shaped by these factors.

Bias↗

Analysis of whole-genome microarray replicates using mixed models.

MOTIVATION: Microarray experiments are inherently noisy. Replication is the key to estimating realistic fold-changes despite such noise. In the analysis of the various sources of noise the dependency structure of the replication needs to be taken into account. RESULTS: We analyzed replicate data sets from a Mycobacterium tuberculosis trcS mutant in order to identify differentially expressed genes and suggest new methods for filtering and normalizing raw array data and for imputing missing values. Mixed ANOVA models are applied to quantify the various sources of error. Such analysis also allows us to determine the optimal number of samples and arrays. Significance values for differential expression are obtained by a hierarchical bootstrapping scheme on scaled residuals. Four highly upregulated genes, including bfrB, were analyzed further. We observed an artefact, where transcriptional readthrough from these genes led to apparent upregulation of adjacent genes. AVAILABILITY: All methods and data discussed are available in the package YASMAhttp://www.cryst.bbk.ac.uk/wernisch/yasma.html for the statistical data analysis system R (http://www.R-project.org).

Algorithms↗

Folding free energy function selects native-like protein sequences in the core but not on the surface.

An automatic protein design procedure is used to select amino acid sequences that optimize the folding free energy function for a given protein. The only information used in designing the sequences is a set of known backbone structures for each protein, a rotamer library, and a well established classical empirical force field, which relies on basic physical chemical principles that underlie molecular interactions and protein stability, and has not been adjusted to yield native-like sequences. Applying the procedure to 7 different known protein folds, representing a total of 45 different native protein structures, yields ensembles of designed sequences displaying remarkable similarity to their natural counterparts in the protein core, but which are distinctly non-native on the protein surface. We show that natural and designed sequences for a given fold score significantly higher than random sequences against profiles derived from both, designed and natural sequence ensembles. Furthermore, we find that designed sequence profiles can be used to retrieve the native sequences for many of the analyzed proteins using standard PSI-BLAST searches in sequence databases. These findings may have important implications for our understanding the selection pressures operating on natural protein sequences and hold promise for improving fold recognition.

Amino Acid Sequence↗

Dissection of the heat-shock response in Mycobacterium tuberculosis using mutants and microarrays.

Regulation of the expression of heat-shock proteins plays an important role in the pathogenesis of Mycobacterium tuberculosis. The heat-shock response of bacteria involves genome-wide changes in gene expression. A combination of targeted mutagenesis and whole-genome expression profiling was used to characterize transcription factors responsible for control of genes encoding the major heat-shock proteins of M. tuberculosis. Two heat-shock regulons were identified. HspR acts as a transcriptional repressor for the members of the Hsp70 (DnaK) regulon, and HrcA similarly regulates the Hsp60 (GroE) response. These two specific repressor circuits overlap with broader transcriptional changes mediated by alternative sigma factors during exposure to high temperatures. Several previously undescribed heat-shock genes were identified as members of the HspR and HrcA regulons. A novel HspR-controlled operon encodes a member of the low-molecular-mass alpha-crystallin family. This protein is one of the most prominent features of the M. tuberculosis heat-shock response and is related to a major antigen induced in response to anaerobic stress.

Bacterial Proteins↗