GenBank: current status and future directions.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to C Burks.
Explore the source record for details and available documents.
The rapidly increasing number of databases relevant to molecular biology has given rise to a need for a coordinated effort to identify, characterize, and link them. The LiMB database, which contains information about molecular biology and related databases, is a step in that direction. It serves molecular biologists seeking data sets containing information relevant to their research, and is also intended to anticipate the needs of database designers and managers building software links for related data sets. We present an abbreviated version of the database here; the full database is available free of charge as described below.
The lipid state hypothesis proposes that liquid crystalline states of cholesteryl esters play a role in the development and persistence of the fatty streak lesions characteristic of atherosclerosis. We have tested several corollaries suggested by this hypothesis and find that the ensemble of droplets in atherosclerotic tissue are predominantly in the isotropic (fluid) state at 37.0 degrees C. Furthermore, the liquid-crystalline state transition behavior of these droplets is not influenced significantly by the distribution of component cholesteryl ester species. There are no significant correlations between the transition behavior of the droplets and the age, sex, or race of the subjects from which tissue samples were taken. These results show that the lipid state hypothesis is weak, and that the origin and persistence of fatty streak lesions in humans is probably dominated by other factors.
The distribution of interspersed repetitive DNA sequences in the human genome has been investigated, using a combination of biochemical, cytological, computational, and recombinant DNA approaches. "Low-resolution" biochemical experiments indicate that the general distribution of repetitive sequences in human DNA can be adequately described by models that assume a random spacing, with an average distance of 3 kb. A detailed "high-resolution" map of the repetitive sequence organization along 400 kb of cloned human DNA, including 150 kb of DNA fragments isolated for this study, is consistent with this general distribution pattern. However, a higher frequency of spacing distances greater than 9.5 kb was observed in this genomic DNA sample. While the overall repetitive sequence distribution is best described by models that assume a random distribution, an analysis of the distribution of Alu repetitive sequences appearing in the GenBank sequence database indicates that there are local domains with varying Alu placement densities. In situ hybridization to human metaphase chromosomes indicates that local density domains for Alu placement can be observed cytologically. Centric heterochromatin regions, in particular, are at least 50-fold underrepresented in Alu sequences. The observed distribution for repetitive sequences in human DNA is the expected result for sequences that transpose throughout the genome, with local regions of "preference" or "exclusion" for integration.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
The GenBank Genetic Sequence Data Bank contains nearly 15,000 entries for DNA and RNA sequences that have been reported since 1967. This paper briefly describes the contents of the database, the forms in which the data are distributed, and the services available to scientists using the GenBank database.
Explore the source record for details and available documents.
We report 3 cases of penile curvature deformity and fibrosis of the tunica albuginea after long-term self-injection with papaverine hydrochloride and phentolamine mesylate. To our knowledge this complication has not been reported previously. The clinical presentation and possible etiology are discussed.
A growing body of data indicates that the equilibrium structures of some DNA fragments are curved and that curvature is sequence-directed. We describe a quantitative measure of DNA curvature that can be used for evaluating and comparing current proposed models for the molecular basis of DNA curvature. We demonstrate that this measure, in conjunction with any given prediction model, enables both the comparison of experimental data to predictions and the scanning of nucleotide sequence databases for potential curved regions.
We have searched the GenBank nucleic acid sequence database for potential short restriction fragments. All possible oligonucleotides up to length five are found at least once flanked by known restriction recognition patterns. Thus, searches in the database for a specific sequence corresponding to a desired oligonucleotide would often point to one or more sources of short, retrievable fragments containing that sequence. These results underscore the potential of nucleic acid sequence databases in planning experiments.
The GenBank Genetic Sequence Data Bank contains over 5700 entries for DNA and RNA sequences that have been reported since 1967. This paper briefly describes the contents of the database, the forms in which the database is distributed, and the services we offer to scientists who use the GenBank database.
All pairs of a large set of known vertebrate DNA sequences were searched by computer for most similar segments. Analysis of this data shows that the computed similarity scores are distributed proportionally to the logarithm of the product of the lengths of the sequences involved. This distribution is closely related to recent results of Erdos and others on the longest run of heads in coin tossing. A simple rule is derived for determination of statistical significance of the similarity scores and to assist in relating statistical and biological significance.
The GenBank nucleic acid sequence database is a computer-based collection of all published DNA and RNA sequences; it contains over five million bases in close to six thousand sequence entries drawn from four thousand five hundred published articles. Each sequence is accompanied by relevant biological annotation. The database is available either on magnetic tape, on floppy diskettes, on-line or in hardcopy form. We discuss the structure of the database, the extent of the data and the implications of the database for research on nucleic acids.
The motA and motB gene products of Escherichia coli are integral membrane proteins necessary for flagellar rotation. We determined the DNA sequence of the region containing the motA gene and its promoter. Within this sequence, there is an open reading frame of 885 nucleotides, which with high probability (98% confidence level) meets criteria for a coding sequence. The 295-residue amino acid translation product had a molecular weight of 31,974, in good agreement with the value determined experimentally by gel electrophoresis. The amino acid sequence, which was quite hydrophobic, was subjected to a theoretical analysis designed to predict membrane-spanning alpha-helical segments of integral membrane proteins; four such hydrophobic helices were predicted by this treatment. Additional amphipathic helices may also be present. A remarkable feature of the sequence is the existence of two segments of high uncompensated charge density, one positive and the other negative. Possible organization of the protein in the membrane is discussed. Asymmetry in the amino acid composition of translated DNA sequences was used to distinguish between two possible initiation codons. The use of this method as a criterion for authentication of coding regions is described briefly in an Appendix.
Explore the source record for details and available documents.