Search PubMed⌕ Search

Biomedical subjects

David B Searls

Publications and source records attributed to David B Searls.

8 recordsLinked to original sources

Grammatical representations of macromolecular structure.

Since the first application of context-free grammars to RNA secondary structures in 1988, many researchers have used both ad hoc and formal methods from computational linguistics to model RNA and protein structure. We show how nearly all of these methods are based on the same core principles and can be converted into equivalent approaches in the framework of tree-adjoining grammars and related formalisms. We also propose some new approaches that extend these core principles in novel ways.

Algorithms↗

Data integration: challenges for drug discovery.

The effective integration of data and knowledge from many disparate sources will be crucial to future drug discovery. Data integration is a key element of conducting scientific investigations with modern platform technologies, managing increasingly complex discovery portfolios and processes, and fully realizing economies of scale in large enterprises. However, viewing data integration as simply an 'IT problem' underestimates the novel and serious scientific and management challenges it embodies - challenges that could require significant methodological and even cultural changes in our approach to data.

Data Collection↗

Clusters of adjacent and similarly expressed genes across normal human tissues complicate comparative transcriptomic discovery.

Transcriptomic techniques are valuable tools with which to validate genetic and biological hypotheses and are now widely available for research. However, with the exception of tumor biology, comparative genomics analyses have been difficult to use as discovery engines to describe biologically relevant expression changes. We propose that physical proximity of human genes correlates with similar mRNA expression, so that increased expression might include a disease-relevant gene and many other genes in the adjacent region. To increase the efficiency of combining susceptibility gene mapping and interpretation of transcriptomics, we developed a method to identify clusters of adjacent and similarly expressed genes. Gene expression profiles for 28,945 genes across 101 normal human tissues were obtained from the Gene Logic BioExpress system. The expression similarity for genes in sliding-windows was measured using average pair-wise Pearson correlation coefficients. We identified 187 clusters (p < 10e-4) of co-regulated genes, including 2648 genes, or 9.1% of all genes considered and termed these "clusters of adjacent and similarly expressed genes" (CASEGs). Genes in 15 (8.2%) of these clusters demonstrate a significant co-expression enrichment (p < 10e-10). This study demonstrates the coordinate expression of neighboring genes and provides a comprehensive view of expression-based compartmentalization of the human genome, which can be overlaid on genetic susceptibility gene maps.

Gene Expression Profiling↗

The language of genes.

Linguistic metaphors have been woven into the fabric of molecular biology since its inception. The determination of the human genome sequence has brought these metaphors to the forefront of the popular imagination, with the natural extension of the notion of DNA as language to that of the genome as the 'book of life'. But do these analogies go deeper and, if so, can the methods developed for analysing languages be applied to molecular biology? In fact, many techniques used in bioinformatics, even if developed independently, may be seen to be grounded in linguistics. Further interweaving of these fields will be instrumental in extending our understanding of the language of life.

Computational Biology↗