Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “GenBank”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

GenBank mining reveals novel insights into Rhizobium phylogeny: Identical 16S rRNA sequences are mainly uncoupled from species designation, host plant, and geographic origin: How this search suggested the definition of a direct 'microbial h-index'.

16S rDNA is the historical gold standard for bacterial identification, particularly in metabarcoding approaches reliant on sequence similarity thresholds. We analyzed 6,660 Rhizobium 16S rRNA gene sequences from GenBank to examine the relationship between sequence identity and three metadata: species name, host plant, and geographic origin. Using an iterative BLAST-based pipeline, we detected 116,069 pairwise matches and assessed concordance among sequences (average length 1,328 bp) sharing 100% identity. For those in which the organism name, host plant and country of isolation were present in the record, surprisingly, 66.59% of identical sequence pairs showed full discordance across all three metadata, while only 1.40% shared the same name, host, and country. The most widespread sequence, detected 371 times, was associated with over 56 different host plants across 25 countries and bore multiple species name designations. These results highlight a striking mismatch between the 16S barcode and the taxonomic, ecological, and phenotypic variability it is assumed to reflect, likely arising from the slow evolution of rRNA genes contrasted with the mobility of ecologically relevant genes via horizontal transfer on plasmids, transposons, and phages. Our findings further challenge the limitations of relying on 16S rRNA alone for fine-scale taxonomic and metadata-based inference in capturing the true functional and ecological diversity of bacteria, endorsing the critical importance of polyphasic taxonomic approaches that integrate genomic, phenotypic, and ecological data. An interesting byproduct of the analysis was to realize the possibility of treating these data as if they were 'citations.' The more one finds the same query sequence, the more that sequence can be considered biologically 'cited', i.e., re-proposed elsewhere in the world. Thus, one can also analyze the h-index of such a ranking. In our Rhizobium dataset, we calculated an h-index = 201, meaning the sequence ranked 201st had 202 identical homologues in GenBank. Although the research effort on given species is directly connected with it, this number provides a quantitative indicator of a taxon's sequence recurrence and distribution within public databases, independent of nomenclatural inconsistencies, offering a novel framework for assessing bacterial representation across global datasets.

RNA, Ribosomal, 16S↗

A novel method to develop highly specific models for regulatory units detects a new LTR in GenBank which contains a functional promoter.

Functional promoters are composed of individual modules (e.g. transcription factor binding sites, secondary structure elements, repeats) arranged in distinct patterns. Recognition of such patterns is essential for identification of promoters in non-coding sequences. However, this is difficult due to the absence of overall sequence similarity in promoters even if they are regulated in a similar way. We implemented simple formal representations of general features of regulatory regions into an algorithm capable of developing complex models reflecting both the element composition and the functional organization of individual elements (ModelGenerator). Though ModelGenerator requires a very simple initial model (e.g. two modules and their relative order) it will generate a much more sophisticated model by analysis of the training set of at least ten sequences. We show ModelGenerator to successfully model different retroviral long terminal repeat (LTR) classes (Lentivirus as well as avian and mammalian C-type) which contain functional promoters. Database searches with the program ModelInspector demonstrated the high specificity of these models and no apparent false negatives were detected. We also verified one match from GenBank to the mammalian C-type LTR model experimentally and showed this sequence to contain an active promoter. Thus, the concept of modular organization of functional regulatory DNA regions (e.g. promoters) could be successfully implemented into a set of computer tools which might be flexible and specific enough to be suitable for prospective analysis of new genomic DNA sequences.

Algorithms↗

Presence of cloning vector sequences in the untranslated region of some genes in Genbank.

We found multiple cloning site sequences in the reported untranslated regions (UTR) of several genes in Genbank. The erroneous information can result in the failure to amplify DNA fragments containing untranslated regions by RT-PCR. It is suggested that a BLAST search be performed when primers are designed for PCR amplification of the 5' or 3' UTR of genes to ensure that the reported UTR does not contain plasmid-derived sequences.

3' Untranslated Regions↗

Bias explorer: measurements of compositional bias in EMBL and GenBank sequence files.

A Windows application for compositional analysis of sequenced genomes (EMBL or GenBank flat files) is available as freeware. The application allows the user to quantify word bias using Markov chain analysis and it allows the user to generate sliding window data for GC-skew, AT-skew, purine excess, keto excess and discrete word counts. The mathematical routines reside in a dynamic link library (DLL), which can be used independently by other applications. The software is available for download at http://www.dfuni.dk/~anfu/Bioinformatics/Main.htm.

Bias↗

Giant G+C% mosaic structures of the human genome found by arrangement of GenBank human DNA sequences according to genetic positions.

To determine the overall variation in the G+C% distribution over long ranges of the human genome, DNA sequences of human genes, which were closely linked genetically or physically, were surveyed from the GenBank Data Bank. A total of 72 sequences longer than 2 kb, which were mutually linked within 500 kb, were identified. The sequences belonged to 17 linkage groups and were ordered in each group according to their genetic positions. Analyses of the G+C% distribution along the ordered sequences showed that sequences within each group almost always had similar G+C% levels, but those belonging to different groups often had different levels. Similar analyses of more distantly linked sequences (e.g., greater than 10 Mb) showed mosaic structures of G+C% distribution. These findings are consistent with predictions made from the "isochore" structures found by CsCl equilibrium centrifugation, in that the structures having homogeneous base compositions stretched over at least several hundred kilobases. A possible boundary of the giant G+C% mosaic structures was identified between X-linked G6PD and F8C.

Base Composition↗

Divergence of Genbank and human tumor Bcl-2 sequences and implications for binding affinity to key apoptotic proteins.

Heterodimerization of antiapoptotic and pro-apoptotic Bcl-2 family of proteins provides an important mechanism for apoptosis regulation. Knowledge about key amino acids in the binding groove of native Bcl-2 contributing to this interaction will greatly facilitate the design of Bcl-2-specific inhibitors. There are two different Bcl-2 sequences, M13994 and M14745, in Genbank. Chimeric proteins Bcl-2(1) and Bcl-2(2) derived from the above sequences, although similar in structure, showed different binding affinities to Bak and Bad BH3 peptides (Petros et al., 2001). In this study, we show that the Bcl-2(1) sequence in normal and tumor human tissue samples differs from M13994 and M14745, and contains P59, T96, R110, S117 and G237. The actual sequence in the binding pocket matches the Bcl-2-Ig fusion sequence X06487, originally identified in a t(14:18) translocation of the Bcl-2 gene, associated with follicular lymphoma. The possible effects of the observed amino acid differences compared to M13994 and M14745 were investigated by combining structural data with fluorescence anisotropy. G110R substitution confers on Bcl-2(1) substantially increased binding affinity to Bak, Bad and Bax BH3 peptides, demonstrating that R110 is a key contributor to the BH3 binding affinity of Bcl-2. Although NMR structure did not predict R110 involvement in binding to these BH3 peptides, fluorescence anisotropy data clearly points to a critical role for this residue in binding to pro-apoptotic Bcl-2 family members.

Amino Acid Sequence↗

A mechanism for maintaining an up-to-date GenBank database via Usenet.

In this paper, we describe an automated system for distributing updates to the GenBank nucleic acid sequence database, using the Usenet news system as the underlying transport mechanism. Our system allows new loci to be distributed as soon as the sequences are available, over existing networks, using existing Usenet software and infrastructure currently available on a wide range of computer systems.

Amino Acid Sequence↗

Cleaning the GenBank Arabidopsis thaliana data set.

Data driven computational biology relies on the large quantities of genomic data stored in international sequence data banks. However, the possibilities are drastically impaired if the stored data is unreliable. During a project aiming to predict splice sites in the dicot Arabidopsis thaliana, we extracted a data set from the A.thaliana entries in GenBank. A number of simple 'sanity' checks, based on the nature of the data, revealed an alarmingly high error rate. More than 15% of the most important entries extracted did contain erroneous information. In addition, a number of entries had directly conflicting assignments of exons and introns, not stemming from alternative splicing. In a few cases the errors are due to mere typographical misprints, which may be corrected by comparison to the original papers, but errors caused by wrong assignments of splice sites from experimental data are the most common. It is proposed that the level of error correction should be increased and that gene structure sanity checks should be incorporated--also at the submitter level--to avoid or reduce the problem in the future. A non-redundant and error corrected subset of the data for A.thaliana is made available through anonymous FTP.

Algorithms↗

Virgil: a database of rich links between GDB and GenBank.

Database interconnection requires the development of links between related objects from different databases. We built a database of links, called Virgil, to manage and distribute rich (documented) links between GDB genes and GenBank human sequences. Virgil contains 18 667 unique links. In addition to a simple Web form for ad-hoc queries, we propose a generic Web interface and a prototype CORBA server for link distribution. Materials described in this paper are available from http://www.infobiogen.fr/services/virgil/home. html

Computer Communication Networks↗

PubCrawler: keeping up comfortably with PubMed and GenBank.

The free PubCrawler web service (http://www.pubcrawler.ie) has been operating for five years and so far has brought literature and sequence updates to over 22 000 users. It provides information on a personalized web page whenever new articles appear in PubMed or when new sequences are found in GenBank that are specific to customized queries. The server also acts as an automatic alerting system by sending out short notifications or emails with the latest updates as soon as they become available. A new output format and more flexibility for the email formatting help PubCrawler cope with increasing challenges arising from browser incompatibilities and mail filters, therefore making it suitable for a wide range of users.

Databases, Nucleic Acid↗

Intraspecific variation in small-subunit rRNA sequences in GenBank: why single sequences may not adequately represent prokaryotic taxa.

Small-subunit rRNA (SSU rRNA) sequencing is a powerful tool to detect, identify, and classify prokaryotic organisms, and there is currently an explosion of SSU rRNA sequencing in the microbiology community. We report unexpectedly high levels of intraspecific variation (within and between strains) of prokaryote SSU rRNA sequences deposited in GenBank. A total of 82% of the prokaryote species with two published SSU rRNA sequences had more variable positions than a 0.1% random sequencing error would predict, and 48% of these sequence pairs had more variable positions than predicted by a 1.0% random sequencing error. Other sources of sequence variability must account for some of this intraspecific variation. Given these results, phylogenetic studies and biodiversity estimates obtained by using prokaryotic SSU rRNA sequences cannot proceed under the assumption that rRNA sequences of single operons from single isolates adequately represent their taxa. Sequencing SSU rRNA molecules from multiple operons and multiple isolates is highly recommended to obtain meaningful phylogenetic hypotheses, as is careful attention to accurate strain identification.

Bacteria↗

Long-range correlation properties of coding and noncoding DNA sequences: GenBank analysis.

An open question in computational molecular biology is whether long-range correlations are present in both coding and noncoding DNA or only in the latter. To answer this question, we consider all 33301 coding and all 29453 noncoding eukaryotic sequences--each of length larger than 512 base pairs (bp)--in the present release of the GenBank to dtermine whether there is any statistically significant distinction in their long-range correlation properties. Standard fast Fourier transform (FFT) analysis indicates that coding sequences have practically no correlations in the range from 10 bp to 100 bp (spectral exponent beta=0.00 +/- 0.04, where the uncertainty is two standard deviations). In contrast, for noncoding sequences, the average value of the spectral exponent beta is positive (0.16 +/- 0.05) which unambiguously shows the presence of long-range correlations. We also separately analyze the 874 coding and the 1157 noncoding sequences that have more than 4096 bp and find a larger region of power-law behavior. We calculate the probability that these two data sets (coding and noncoding) were drawn from the same distribution and we find that it is less than 10(-10). We obtain independent confirmation of these findings using the method of detrended fluctuation analysis (DFA), which is designed to treat sequences with statistical heterogeneity, such as DNA's known mosaic structure ("patchiness") arising from the nonstationarity of nucleotide concentration. The near-perfect agreement between the two independent analysis methods, FFT and DFA, increases the confidence in the reliability of our conclusion.

Animals↗

Bovine and ovine DNA microsatellites from the EMBL and GENBANK databases.

Bovine and ovine microsatellite sequences were extracted from the EMBL and GENBANK databases. When analysed for number of alleles and degree of heterozygosity in the CSIRO cattle reference families, allele numbers range from 1 to 14 with heterozygosities, in the polymorphic systems ranging from 15.8% to 100%. Six (46%) of the 13 bovine systems tested gave specific and polymorphic products in sheep. Similarly 2 of the 4 ovine systems gave specific and polymorphic products in cattle. These data define 11 bovine and 8 ovine microsatellite systems which are associated with known genes and are thus useful for comparative mapping studies.

Alleles↗

Electronic data publishing and GenBank.

GenBank, the national repository for nucleotide sequence data, has implemented a new model of scientific data management, which we term electronic data publishing. In traditional publishing, both scientific conclusions and supporting data are communicated via the printed page, and in electronic journal publishing, both types of information are communicated via electronic media. In electronic data publishing, by contrast, conclusions are published in a journal while data are published via a network-accessible, electronic database.

Base Sequence↗

Biomedical database inter-connectivity: an experiment linking MIM, GENBANK, and META-1 via MEDLINE.

The linkage of disparate biomedical databases is an important goal of the Unified Medical Language (UMLS) Project. We conducted an experiment to investigate the feasibility of using UMLS resources to link databases in clinical genetics and molecular biology. References from MIM ("Mendelian Inheritance in Man") were lexically mapped to the equivalent citations in MEDLINE. The MeSH major subject headings by which the citations in a particular MIM entry had been indexed were used to develop a "genetic-disorder-centered view of the world" in Meta-1 (the first official version of the UMLS Metathesaurus). Our hypothesis was that these MeSH subject headings could provide access to a "semantic neighborhood" in Meta-1 that would be relevant to a particular genetic disorder. By browsing in this "semantic neighborhood," a user could select various combinations of terms with which to search MEDLINE through an interface between Meta-1 and Grateful Med. Such searches might retrieve citations that were more recent than those in MIM or that provided useful supplementary information. Since some MEDLINE records contain pointers to entries in GENBANK, information about genetic sequences related to a particular clinical genetic disorder could also be retrieved. This scenario was implemented for a small number of MIM entries, providing a concrete demonstration that linking disparate electronic databases in an important subdomain of biomedicine is relatively straightforward.

Databases, Factual↗

Reorganization and merging of the EMBL and GenBank keyword indexes in a tree structure for more efficient retrieval of nucleic acid sequences.

EMBL and GenBank keyword indexes have no hierarchical structure. In this paper we present a method for merging and reorganizing them in a tree structure whose primary roots are the keywords 'protein', 'DNA', 'RNA', and 'unclassified'. Synonymous keywords have been grouped together and erroneous keywords have been corrected. This taxonomic organization of keywords results in a more extensive and efficient retrieval which is further aided by "synonyms declaration". The tree has been produced using the computer programs GENPOINT and CREANET.

Abstracting and Indexing↗

The GenBank nucleic acid sequence database.

The GenBank nucleic acid sequence database is a computer-based collection of all published DNA and RNA sequences; it contains over five million bases in close to six thousand sequence entries drawn from four thousand five hundred published articles. Each sequence is accompanied by relevant biological annotation. The database is available either on magnetic tape, on floppy diskettes, on-line or in hardcopy form. We discuss the structure of the database, the extent of the data and the implications of the database for research on nucleic acids.

Base Sequence↗