Search PubMedSearch

Biomedical subjects

C Burks

Publications and source records attributed to C Burks.

At least 19 recordsLinked to original sources

Splicing signals in Drosophila: intron size, information content, and consensus sequences.

A database of 209 Drosophila introns was extracted from Genbank (release number 64.0) and examined by a number of methods in order to characterize features that might serve as signals for messenger RNA splicing. A tight distribution of sizes was observed: while the smallest introns in the database are 51 nucleotides, more than half are less than 80 nucleotides in length, and most of these have lengths in the range of 59-67 nucleotides. Drosophila splice sites found in large and small introns differ in only minor ways from each other and from those found in vertebrate introns. However, larger introns have greater pyrimidine-richness in the region between 11 and 21 nucleotides upstream of 3' splice sites. The Drosophila branchpoint consensus matrix resembles C T A A T (in which branch formation occurs at the underlined A), and differs from the corresponding mammalian signal in the absence of G at the position immediately preceding the branchpoint. The distribution of occurrences of this sequence suggests a minimum distance between 5' splice sites and branchpoints of about 38 nucleotides, and a minimum distance between 3' splice sites and branchpoints of 15 nucleotides. The methods we have used detect no information in exon sequences other than in the few nucleotides immediately adjacent to the splice sites. However, Drosophila resembles many other species in that there is a discontinuity in A + T content between exons and introns, which are A + T rich.

Animals

GenBank.

The GenBank nucleotide sequence database now contains sequence data and associated annotation corresponding to 85,000,000 nucleotides in 67,000 entries from a total of 3,000 organisms. The input stream of data coming into the database is primarily as direct submissions from the scientific community on electronic media, with little or no data being keyboarded from the printed page by the databank staff. The data are maintained in a relational database management system and are made available in flatfile form through on-line access, and through various network and off-line computer-readable media. The data are also distributed in relational form through satellite copies at a number of institutions in the U.S. and elsewhere. In addition, GenBank provides the U.S. distribution center for the BIOSCI electronic bulletin board service.

Animals

Identifying potential tRNA genes in genomic DNA sequences.

We have developed an algorithm that automatically and reproducibly identifies potential tRNA genes in genomic DNA sequences, and we present a general strategy for testing the sensitivity of such algorithms. This algorithm is useful for the flagging and characterization of long genomic sequences that have not been experimentally analyzed for identification of functional regions, and for the scanning of nucleotide sequence databases for errors in the sequences and the functional assignments associated with them. In an exhaustive scan of the GenBank database, 97.5% of the 744 known tRNA genes were correctly identified (true-positives), and 42 previously unidentified sequences were predicted to be tRNAs. A detailed analysis of these latter predictions reveals that 16 of the 42 are very similar to known tRNA genes, and we predict that they do, in fact, code for tRNA, yielding a false-positive rate for the algorithm of 0.003%. The new algorithm and testing strategy are a considerable improvement over any previously described strategies for recognizing tRNA genes, and they allow detections of genes (including introns) embedded in long genomic sequences.

Algorithms

Electronic data publishing and GenBank.

GenBank, the national repository for nucleotide sequence data, has implemented a new model of scientific data management, which we term electronic data publishing. In traditional publishing, both scientific conclusions and supporting data are communicated via the printed page, and in electronic journal publishing, both types of information are communicated via electronic media. In electronic data publishing, by contrast, conclusions are published in a journal while data are published via a network-accessible, electronic database.

Base Sequence

GenBank.

The GenBank nucleotide sequence database now contains sequence data and associated annotation corresponding to 56,000,000 nucleotides in 45,000 entries. The input stream of data coming into the database has largely been shifted to direct submissions from the scientific community on electronic media. The data have been installed in a relational database management system and are made available in this form through on-line access, and through various network and off-line computer-readable media. In addition, GenBank provides the U.S. distribution center for the BIOSCI electronic bulletin board service.

Base Sequence

A program for computer-assisted scoring of Southern blots.

SCORE, a program for computer-assisted scoring of Southern blots of clone DNA, retains the use of expert human judgment while taking over much of the drudgery of the scoring task. The primary functions of the program are to help make an aligned overlay of the fluorescence gel image and the autoradiogram blot image, to keep track of band and lane locations and to store the resulting data directly into a database. Use of SCORE has resulted in greatly increased efficiency and accuracy.

Autoradiography

Overview of the LiMB database.

The rapidly increasing number of databases relevant to molecular biology has given rise to a need for a coordinated effort to identify, characterize, and link them. The LiMB database, which contains information about molecular biology and related databases, is a step in that direction. It serves molecular biologists seeking data sets containing information relevant to their research, and is also intended to anticipate the needs of database designers and managers building software links for related data sets. We present an abbreviated version of the database here; the full database is available free of charge as described below.

Information Systems

Limitations of the lipid state hypothesis for atherosclerosis are revealed by X-ray diffraction measurements.

The lipid state hypothesis proposes that liquid crystalline states of cholesteryl esters play a role in the development and persistence of the fatty streak lesions characteristic of atherosclerosis. We have tested several corollaries suggested by this hypothesis and find that the ensemble of droplets in atherosclerotic tissue are predominantly in the isotropic (fluid) state at 37.0 degrees C. Furthermore, the liquid-crystalline state transition behavior of these droplets is not influenced significantly by the distribution of component cholesteryl ester species. There are no significant correlations between the transition behavior of the droplets and the age, sex, or race of the subjects from which tissue samples were taken. These results show that the lipid state hypothesis is weak, and that the origin and persistence of fatty streak lesions in humans is probably dominated by other factors.

Adolescent

The distribution of interspersed repetitive DNA sequences in the human genome.

The distribution of interspersed repetitive DNA sequences in the human genome has been investigated, using a combination of biochemical, cytological, computational, and recombinant DNA approaches. "Low-resolution" biochemical experiments indicate that the general distribution of repetitive sequences in human DNA can be adequately described by models that assume a random spacing, with an average distance of 3 kb. A detailed "high-resolution" map of the repetitive sequence organization along 400 kb of cloned human DNA, including 150 kb of DNA fragments isolated for this study, is consistent with this general distribution pattern. However, a higher frequency of spacing distances greater than 9.5 kb was observed in this genomic DNA sample. While the overall repetitive sequence distribution is best described by models that assume a random distribution, an analysis of the distribution of Alu repetitive sequences appearing in the GenBank sequence database indicates that there are local domains with varying Alu placement densities. In situ hybridization to human metaphase chromosomes indicates that local density domains for Alu placement can be observed cytologically. Centric heterochromatin regions, in particular, are at least 50-fold underrepresented in Alu sequences. The observed distribution for repetitive sequences in human DNA is the expected result for sequences that transpose throughout the genome, with local regions of "preference" or "exclusion" for integration.

Chromosome Mapping

The LiMB database.

Explore the source record for details and available documents.

Information Systems

The GenBank genetic sequence data bank.

The GenBank Genetic Sequence Data Bank contains nearly 15,000 entries for DNA and RNA sequences that have been reported since 1967. This paper briefly describes the contents of the database, the forms in which the data are distributed, and the services available to scientists using the GenBank database.

Base Sequence

A quantitative measure of DNA curvature enabling the comparison of predicted structures.

A growing body of data indicates that the equilibrium structures of some DNA fragments are curved and that curvature is sequence-directed. We describe a quantitative measure of DNA curvature that can be used for evaluating and comparing current proposed models for the molecular basis of DNA curvature. We demonstrate that this measure, in conjunction with any given prediction model, enables both the comparison of experimental data to predictions and the scanning of nucleotide sequence databases for potential curved regions.

Bacteriophage lambda