Search PubMedSearch

Biomedical subjects

D B Searls

Publications and source records attributed to D B Searls.

17 recordsLinked to original sources

Analysis of EST-driven gene annotation in human genomic sequence.

We have performed a systematic analysis of gene identification in genomic sequence by similarity search against expressed sequence tags (ESTs) to assess the suitability of this method for automated annotation of the human genome. A BLAST-based strategy was constructed to examine the potential of this approach, and was applied to test sets containing all human genomic sequences longer than 5 kb in public databases, plus 300 kb of exhaustively characterized benchmark sequence. At high stringency, 70%-90% of all annotated genes are detected by near-identity to EST sequence; >95% of ESTs aligning with well-annotated sequences overlap a gene. These ESTs provide immediate access to the corresponding cDNA clones for follow-up laboratory verification and subsequent biologic analysis. At lower stringency, up to 97% of annotated genes were identified by similarity to ESTs. The apparent false-positive rate rose to 55% of ESTs among all sequences and 20% among benchmark sequences at the lowest stringency, indicating that many genes in public database entries are unannotated. Approximately half of the alignments span multiple exons, and thus aid in the construction of gene predictions and elucidation of alternative splicing. In addition, ESTs from multiple cDNA libraries frequently cluster over genes, providing a starting point for crude expression profiles. Clone IDs may be used to form EST pairs, and particularly to extend models by associating alignments of lower stringency with high-quality alignments. These results demonstrate that EST similarity search is a practical general-purpose annotation technique that complements pattern recognition methods as a tool for gene characterization.

Base Sequence

Computational gene discovery and human disease.

Bioinformatics is now an essential tool in many aspects of human molecular genetics research. Methods for the prediction of gene structure are essential components in genomic sequencing projects and provide the key to deriving protein sequence and locating intron/exon junctions. Sequence comparison and database searching are the pre-eminent approaches for predicting the likely biochemical function of new genes, although sequence profiles derived from families of aligned sequences have advantages in the detection of remote sequence relationships. The use of sequence database analysis for large-scale comparative analysis of genome sequence data from model organisms is emerging as the most important recent development in the application of bioinformatics methods for characterizing candidate disease genes.

Animals

Linguistic approaches to biological sequences.

Biologists have long made use of linguistic metaphors in describing and naming cellular processes involving nucleic acid and protein sequences. Indeed, it is very natural to view the genetic 'text' and its sequential transliterations in these terms. However, a metaphor is not a tool, and it is necessary to ask whether the techniques used in analyzing other kinds of languages, such as human and computer languages, can in fact be of any use in tackling problems in molecular biology. This paper reviews the work of the author and others in applying the methods of computational linguistics to biological sequences.

Base Sequence

bioTk:componentry for genome informatics graphical user interfaces.

bioTk is a collection of graphical "widgets" and utilities that support application programming in the domain of bioinformatics. It is intended to establish a framework that encourages the development of communicating window-based applications and flexible, non-modal user interaction. The current release of bioTk has domain-specific widgets for chromosome ideogram displays, genome maps, and scrolling sequence windows.

Base Sequence

Automata-theoretic models of mutation and alignment.

Finite-state automata called transducers, which have both input and output, can be used to model simple mechanisms of biological mutation. We present a methodology whereby numerically-weighted versions of such specifications can be mechanically adapted to create string edit machines that are essentially equivalent to recurrence relations of the sort that characterize dynamic programming alignment algorithms. Based on this, we have developed a visual programming system for designing new alignment algorithms in a rapid-prototyping fashion.

Algorithms

Gene structure prediction by linguistic methods.

The higher-order structure of genes and other features of biological sequences can be described by means of formal grammars. These grammars can then be used by general-purpose parsers to detect and to assemble such structures by means of syntactic pattern recognition. We describe a grammar and parser for eukaryotic protein-encoding genes, which by some measures is as effective as current connectionist and combinatorial algorithms in predicting gene structures for sequence database entries. Parameters of the grammar rules are optimized for several different species, and mixing experiments are performed to determine the degree of species specificity and the relative importance of compositional, signal-based, and syntactic components in gene prediction.

Animals

SORTEZ: a relational translator for NCBI's ASN.1 database.

The National Center for Biotechnology Information (NCBI) has created a database collection that includes several protein and nucleic acid sequence databases, a biosequence-specific subset of MEDLINE, as well as value-added information such as links between similar sequences. Information in the NCBI database is modeled in Abstract Syntax Notation 1 (ASN.1) an Open Systems Interconnection protocol designed for the purpose of exchanging structured data between software applications rather than as a data model for database systems. While the NCBI database is distributed with an easy-to-use information retrieval system, ENTREZ, the ASN.1 data model currently lacks an ad hoc query language for general-purpose data access. For that reason, we have developed a software package, SORTEZ, that transforms the ASN.1 database (or other databases with nested data structures) to a relational data model and subsequently to a relational database management system (Sybase) where information can be accessed through the relational query language, SQL. Because the need to transform data from one data model and schema to another arises naturally in several important contexts, including efficient execution of specific applications, access to multiple databases and adaptation to database evolution this work also serves as a practical study of the issues involved in the various stages of database transformation. We show that transformation from the ASN.1 data model to a relational data model can be largely automated, but that schema transformation and data conversion require considerable domain expertise and would greatly benefit from additional support tools.

Algorithms

Doing sequence analysis with your printer.

The software package RSVP (Rapid Sequence Visualization in PostScript) has a suite of visually oriented sequence analysis routines implemented entirely in the page description language PostScript, a widely used standard that is built into many printers. RSVP is thus a relatively platform-independent tool for providing a 'quick look' at sequence data, using form and color to help point out patterns, in advance of more sophisticated sequence analyses.

Amino Acid Sequence

Fast Fourier transform-based correlation of DNA sequences using complex plane encoding.

The detection of similarities between DNA sequences can be accomplished using the signal-processing technique of cross-correlation. An early method used the fast Fourier transform (FFT) to perform correlations on DNA sequences in O(n log n) time for any length sequence. However, this method requires many FFTs (nine), runs no faster if one sequence is much shorter than the other, and measures only global similarity, so that significant short local matches may be missed. We report that, through the use of alternative encodings of the DNA sequence in the complex plane, the number of FFTs performed can be traded off against (i) signal-to-noise ratio, and (ii) a certain degree of filtering for local similarity via k-tuple correlation. Also, when comparing probe sequences against much longer targets, the algorithm can be sped up by decomposing the target and performing multiple small FFTs in an overlap-save arrangement. Finally, by decomposing the probe sequence as well, the detection of local similarities can be further enhanced. With current advances in extremely fast hardware implementations of signal-processing operations, this approach may prove more practical than heretofore.

Algorithms

Chromosomal site of hepatitis B virus (HBV) integration in a human hepatocellular carcinoma-derived cell line.

The single site of integration of hepatitis B virus in the human hepatocellular carcinoma cell line Hep 3B 2-1/7 was found to segregate with human chromosome 12 in somatic cell hybrids. Analysis of metaphase spreads of Hep 3B 2-1/7 following in situ hybridization with pHBV revealed integration at 12q13----q14, a location that coincides with a fragile site, fra (12q13). The possible significance of this location to the development of hepatocellular carcinomas is discussed.

Animals

Analysis of early antigenic changes on heterokaryons between L-cells and a teratocarcinoma-derived cell line, TerC.

A method is presented for the detection of antigenic changes in heterokaryons during the first 24 h after fusion, as detected by antibody- and complement-mediated inhibition of 125IDU uptake. The method employs a variation of the HAT selection protocol, and a mathematical model which allows the reactions of a minority population of heterokaryons to be distinguished from those of the parental cells. This procedure is applied to fusions of clones of the teratocarcinoma-derived cell line TerC, which demonstrate very low levels of H-2b expression, with an L-cell derivative rich in H-2k antigens. Heterokaryons demonstrate an initial tendency for levels of H-2b and H-2k expression to converge, beginning within a few hours of fusion; however, H-2b expression never attains the levels of H-2k, even in long-term hybrids. These results are largely confirmed by quantitative absorptions, except that no diminution of H-2k expression is observed, suggesting that fusion with the TerC parents may transfer a resistance to complement damage to the L-cell parent.

Animals

H-2 expression on a teratocarcinoma-derived cell line, TerC.

The murine teratocarcinoma-derived cell line TerC, previously thought to lack products of the major histocompatibility locus H-2, was shown to express low levels of K- and D-end H-2 specificities with the use of a sensitive cytotoxic assay. The assay, based on the ability of antibody and complement to inhibit uptake of the thymidine analogue [125I]5-iodo-2'-deoxyuridine, detected appropriate public and private specificities, as demonstrated with the use of oligospecific antisera and by absorption analyses. A series of clone of TerC varied only slightly in H-2 expression, and there was no tendency for expression to increase with time in culture; thus the low levels of H-2 were not the result of a differentiating subpopulation.

Animals

Lipid composition and lateral diffusion in plasma membranes of teratocarcinoma-derived cell lines.

We measured lipid lateral diffusion rates for a series of teratocarcinoma-derived and embryo-derived cell lines, using the technique of fluorescence photobleaching recovery with a fluorescent lipid probe, C16dil. The probe diffuses more rapidly in plasma membranes of embryonal carcinoma cells than in plasma membranes of teratocarcinoma-derived endodermal cell lines. When embryonal carcinoma cells are induced to differentiate by treatment with retinoic acid, diffusion constants of C16dil are reduced to levels typical of endoderm. These changes are paralleled by differences in membrane cholesterol content; membrane free cholesterol levels in embryonal carcinoma lines are approximately half those found in endodermal lines, and are markedly increased upon retinoic-acid-induced differentiation.

Animals

Plasma membrane isolation on DEAE-Sephadex beads.

A much-simplified method for the purification of plasma membranes of cultured cells is presented, based upon the attachment of viable cells to nitrocellulose-treated DEAE-Sephadex beads, and their subsequent shearing by hypotonic lysis, agitation on a vortex mixer and sonication. The method is suggested by an older procedure involving attachment to poly-(L-lysine)-coated glass or polyacrylamide beads; the preparation involved in the present method, however, is considerably easier, more rapid and less expensive. Recovery of L-cell plasma membrane marker enzyme activities is approx. 25%, while contamination by internal membrane markers is much less than 1%.

Animals