Search PubMed⌕ Search

Biomedical subjects

Erich Bornberg-Bauer

Publications and source records attributed to Erich Bornberg-Bauer.

5 recordsLinked to original sources

Expression of De Novo Open Reading Frames in Natural Populations of Drosophila melanogaster.

De novo genes, which originate from noncoding DNA, are known to have a high rate of turnover over short evolutionary timescales, such as within a species. Thus, their expression is often lineage- or genetic background-specific. However, little is known about their levels and breadth of expression as populations of a species diverge. In this study, we utilized publicly available RNA-seq data to examine the expression of newly evolved open reading frames (neORFs) in comparison to non- and protein-coding genes in Drosophila melanogaster populations from the derived species range in Europe and the ancestral range in sub-Saharan Africa. Our datasets included two adult tissue types as well as whole bodies at two temperatures for both sexes and three larval/prepupal developmental stages in a single tissue and sex, which allowed us to examine neORF expression and divergence across multiple sample types as well as sex and population. We detected a relatively large proportion (approximately 50%) of annotated neORFs as expressed in the population samples, with neORFs often showing greater expression divergence between populations than non- or protein-coding genes. However, differential expression of neORFs between populations tended to occur in a sample type-specific manner. On the other hand, neORFs displayed less sex-biased expression than the other two gene classes, with the majority of sex-biased neORFs detected in whole bodies, which may be attributable to the presence of the gonads. We also found that neORFs shared among multiple lines in the original set of inbred lines in which they were first detected were more likely to be both expressed and differentially expressed in the new population samples, suggesting that neORFs at a higher frequency (i.e. present in more individuals) within a species are more likely to be functional.

Animals↗

Recombinatoric exploration of novel folded structures: a heteropolymer-based model of protein evolutionary landscapes.

The role of recombination in evolution is compared with that of point mutations (substitutions) in the context of a simple, polymer physics-based model mapping between sequence (genotype) and conformational (phenotype) spaces. Crossovers and point mutations of lattice chains with a hydrophobic polar code are investigated. Sequences encoding for a single ground-state conformation are considered viable and used as model proteins. Point mutations lead to diffusive walks on the evolutionary landscape, whereas crossovers can "tunnel" through barriers of diminished fitness. The degree to which crossovers allow for more efficient sequence and structural exploration depends on the relative rates of point mutations versus that of crossovers and the dispersion in fitness that characterizes the ruggedness of the evolutionary landscape. The probability that a crossover between a pair of viable sequences results in viable sequences is an order of magnitude higher than random, implying that a sequence's overall propensity to encode uniquely is embodied partially in local signals. Consistent with this observation, certain hydrophobicity patterns are significantly more favored than others among fragments (i.e., subsequences) of sequences that encode uniquely, and examples reminiscent of autonomous folding units in real proteins are found. The number of structures explored by both crossovers and point mutations is always substantially larger than that via point mutations alone, but the corresponding numbers of sequences explored can be comparable when the evolutionary landscape is rugged. Efficient structural exploration requires intermediate nonextreme ratios between point-mutation and crossover rates.

Biophysical Phenomena↗

Conceptual data modelling for bioinformatics.

Current research in the biosciences depends heavily on the effective exploitation of huge amounts of data. These are in disparate formats, remotely dispersed, and based on the different vocabularies of various disciplines. Furthermore, data are often stored or distributed using formats that leave implicit many important features relating to the structure and semantics of the data. Conceptual data modelling involves the development of implementation-independent models that capture and make explicit the principal structural properties of data. Entities such as a biopolymer or a reaction, and their relations, eg catalyses, can be formalised using a conceptual data model. Conceptual models are implementation-independent and can be transformed in systematic ways for implementation using different platforms, eg traditional database management systems. This paper describes the basics of the most widely used conceptual modelling notations, the ER (entity-relationship) model and the class diagrams of the UML (unified modelling language), and illustrates their use through several examples from bioinformatics. In particular, models are presented for protein structures and motifs, and for genomic sequences.

Computational Biology↗

TreeWiz: interactive exploration of huge trees.

MOTIVATION: The rapidly increasing amount and disparity of biological data requires interpretation at many levels of description. Human judgement and intuition are important because not all data can be automatically and comprehensively analyzed. Visualization of trees and substructures corresponding to certain features are often used to analyze phylogenies or taxonomies. Unfortunately, most existing tools do not cope with the size of current datasets, the required functionality, or both. RESULTS: We introduce a program for visualization of huge trees and also for the interactive exploration of their content. We have developed a range of new schemes which are tailored for biological problems. Users can get an overview, zoom in, filter out data and retrieve details from standard databases such as SWISS-PROT. Furthermore, it is possible to analyze the relationship between chosen leaf sets that are specified by common features on a second level of representation. On a PC (with approximately equal to 512 MB RAM), trees of up to several tens of thousands of leaves can be loaded and both rapidly and interactively explored. We demonstrate the use of this program for the analysis of the SYSTERS data set (which contains hierarchically clustered protein sequences) to which PFAM domains were added as features.

Classification↗