Search PubMed⌕ Search

Biomedical subjects

Bernhard Haubold

Publications and source records attributed to Bernhard Haubold.

5 recordsLinked to original sources

Genome comparison without alignment using shortest unique substrings.

BACKGROUND: Sequence comparison by alignment is a fundamental tool of molecular biology. In this paper we show how a number of sequence comparison tasks, including the detection of unique genomic regions, can be accomplished efficiently without an alignment step. Our procedure for nucleotide sequence comparison is based on shortest unique substrings. These are substrings which occur only once within the sequence or set of sequences analysed and which cannot be further reduced in length without losing the property of uniqueness. Such substrings can be detected using generalized suffix trees. RESULTS: We find that the shortest unique substrings in Caenorhabditis elegans, human and mouse are no longer than 11 bp in the autosomes of these organisms. In mouse and human these unique substrings are significantly clustered in upstream regions of known genes. Moreover, the probability of finding such short unique substrings in the genomes of human or mouse by chance is extremely small. We derive an analytical expression for the null distribution of shortest unique substrings, given the GC-content of the query sequences. Furthermore, we apply our method to rapidly detect unique genomic regions in the genome of Staphylococcus aureus strain MSSA476 compared to four other staphylococcal genomes. CONCLUSION: We combine a method to rapidly search for shortest unique substrings in DNA sequences and a derivation of their null distribution. We show that unique regions in an arbitrary sample of genomes can be efficiently detected with this method. The corresponding programs shustring (SHortest Unique subSTRING) and shulen are written in C and available at http://adenine.biz.fh-weihenstephan.de/shustring/.

Algorithms↗

Comparative genomics: methods and applications.

Interpreting the functional content of a given genomic sequence is one of the central challenges of biology today. Perhaps the most promising approach to this problem is based on the comparative method of classic biology in the modern guise of sequence comparison. For instance, protein-coding regions tend to be conserved between species. Hence, a simple method for distinguishing a functional exon from the chance absence of stop codons is to investigate its homologue from closely related species. Predicting regulatory elements is even more difficult than exon prediction, but again, comparisons pinpointing conserved sequence motifs upstream of translation start sites are helping to unravel gene regulatory networks. In addition to interspecific studies, intraspecific sequence comparison yields insights into the evolutionary forces that have acted on a species in the past. Of particular interest here is the identification of selection events such as selective sweeps. Both intra- and interspecific sequence comparisons are based on a variety of computational methods, including alignment, phylogenetic reconstruction, and coalescent theory. This article surveys the biology and the central computational ideas applied in recent comparative genomics projects. We argue that the most fruitful method of understanding the functional content of genomes is to study them in the context of related genomic sequences. In particular, such a study may reveal selection, a fundamental pointer to biological relevance.

Animals↗

Identification of farnesoid X receptor beta as a novel mammalian nuclear receptor sensing lanosterol.

Nuclear receptors are ligand-modulated transcription factors. On the basis of the completed human genome sequence, this family was thought to contain 48 functional members. However, by mining human and mouse genomic sequences, we identified FXRbeta as a novel family member. It is a functional receptor in mice, rats, rabbits, and dogs but constitutes a pseudogene in humans and primates. Murine FXRbeta is widely coexpressed with FXR in embryonic and adult tissues. It heterodimerizes with RXRalpha and stimulates transcription through specific DNA response elements upon addition of 9-cis-retinoic acid. Finally, we identified lanosterol as a candidate endogenous ligand that induces coactivator recruitment and transcriptional activation by mFXRbeta. Lanosterol is an intermediate of cholesterol biosynthesis, which suggests a direct role in the control of cholesterol biosynthesis in nonprimates. The identification of FXRbeta as a novel functional receptor in nonprimate animals sheds new light on the species differences in cholesterol metabolism and has strong implications for the interpretation of genetic and pharmacological studies of FXR-directed physiologies and drug discovery programs.

Amino Acid Sequence↗

Calculating the SNP-effective sample size from an alignment.

MOTIVATION: The number of Single Nucleotide Polymorphisms (SNPs) detectable in an alignment is a function of the length and the number of the aligned sequences. The latter is called sample size. However, a typical alignment, for instance obtained as a BLAST-search result of a query sequence against an EST database, does not evenly cover the query sequence. Therefore, it is usually not clear what the actual sample size is. RESULTS: We present a method to calculate the effective sample size, called n(eff), for a given BLAST alignment. This method takes into account that multiple coverage contributes only logarithmically to the SNP yield of a given sequence stretch. We show that the effective sample size n(eff) is usually much smaller than would be expected for a given amount of coverage and illustrate this with two typical examples.

Algorithms↗

Recombination and gene conversion in a 170-kb genomic region of Arabidopsis thaliana.

Arabidopsis thaliana is a highly selfing plant that nevertheless appears to undergo substantial recombination. To reconcile its selfing habit with the observations of recombination, we have sampled the genetic diversity of A. thaliana at 14 loci of approximately 500 bp each, spread across 170 kb of genomic sequence centered on a QTL for resistance to herbivory. A total of 170 of the 6321 nucleotides surveyed were polymorphic, with 169 being biallelic. The mean silent genetic diversity (pi(s)) varied between 0.001 and 0.03. Pairwise linkage disequilibria between the polymorphisms were negatively correlated with distance, although this effect vanished when only pairs of polymorphisms with four haplotypes were included in the analysis. The absence of a consistent negative correlation between distance and linkage disequilibrium indicated that gene conversion might have played an important role in distributing genetic diversity throughout the region. We tested this by coalescent simulations and estimate that up to 90% of recombination is due to gene conversion.

Arabidopsis↗