Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “GenBank”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

ColorHOR--novel graphical algorithm for fast scan of alpha satellite higher-order repeats and HOR annotation for GenBank sequence of human genome.

MOTIVATION: GenBank data are at present lacking alpha satellite higher-order repeat (HOR) annotation. Furthermore, exact HOR consensus lengths have not been reported so far. Given the fast growth of sequence databases in the centromeric region, it is of increasing interest to have efficient tools for computational identification and analysis of HORs from known sequences. RESULTS: We develop a graphical user interface method, ColorHOR, for fast computational identification of HORs in a given genomic sequence, without requiring a priori information on the composition of the genomic sequence. ColorHOR is based on an extension of the key-string algorithm and provides a color representation of the order and orientation of HORs. For the key string, we use a robust 6 bp string from a consensus alpha satellite and its representative nature is tested. ColorHOR algorithm provides a direct visual identification of HORs (direct and/or reverse complement). In more detail, we first illustrate the ColorHOR results for human chromosome 1. Using ColorHOR we determine for the first time the HOR annotation of the GenBank sequence of the whole human genome. In addition to some HORs, corresponding to those determined previously biochemically, we find new HORs in chromosomes 4, 8, 9, 10, 11 and 19. For the first time, we determine exact consensus lengths of HORs in 10 chromosomes. We propose that the HOR assignment obtained by using ColorHOR be included into the GenBank database.

Algorithms↗

GenBank.

The GenBank sequence database continues to expand its data coverage, quality control, annotation content and retrieval services for the scientific community. Besides handling direct submissions of sequence data from authors, GenBank also incorporates DNA sequences from all available public sources; an integrated retrieval system, known as Entrez, also makes available data from the major protein sequence and structural databases, and from U.S. and European patents. MIDLINE abstracts from published articles describing the sequences are also included as an additional source of biological annotation for sequence entries. GenBank supports distribution of the data via FTP, CD-ROM, and E-mail servers. Network server-client programs provide access to an integrated database for literature retrieval and sequence similarity searching.

Amino Acid Sequence↗

Knowledge discovery in GenBank.

We describe various methods designed to discover knowledge in the GenBank nucleic acid sequence database. Using a grammatical model of gene structure, we create a parse tree of a gene using features listed in the FEATURE TABLE. The parse tree infers features that are not explicitly listed, but which follow from the listed features. This method discovers 30% more introns and 40% more exons when applied to a globin gene subset of GenBank. Parse tree construction also entails resolving ambiguity and inconsistency within a FEATURE TABLE. We transform the parse tree into an augmented FEATURE TABLE that represents inferred gene structure explicitly and unambiguously, thereby greatly improving the utility of the FEATURE TABLE to researchers. We then describe various analogical reasoning techniques designed to exploit the homologous nature of genes. We build a classification hierarchy that reflects the evolutionary relationship between genes. Descriptive grammars of gene classes are then induced from the instance grammars of genes. Case based reasoning techniques use these abstract gene class descriptions to predict the presence and location of regulatory features not listed in the FEATURE TABLE. A cross-validation test shows a success rate of 87% on a globin gene subset of GenBank.

Algorithms↗

Cloning of alkaline sphingomyelinase from rat intestinal mucosa and adjusting of the hypothetical protein XP_221184 in GenBank.

Intestinal alkaline sphingomyelinase (alk-SMase) digests sphingomyelin and the process may influence colonic tumorigenesis and cholesterol absorption. We recently identified the gene of human alk-SMase and cloned the cDNA. Cross-species screening of homology in GenBank found a hypothetical rat protein, XP_221184, with 491 amino acid residues, which shares 73% identity with human alk-SMase. Based on the cDNA sequence of this protein, we cloned a cDNA from rat intestinal mucosa by RT-PCR. The cloned cDNA encodes a protein with 439 amino acid residues and higher (85%) identity with human alk-SMase. The cloned cDNA differed from the XP_221184 cDNA in splice sites linking exons 2 and 3, and exons 3 and 4, respectively. In the sequence of the cloned protein, the predicted activity motif, sphingomyelin binding sites, and potential glycosylation sites in human alk-SMase are all conserved. To confirm the cloned protein is the real form of alk-SMase, native alk-SMase was purified from rat intestine and subjected to proteolytic digestion followed by matrix-assisted laser desorption/ionization (MALDI) mass spectrometry and electrospray ionization (ESI) tandem mass spectrometry. Seven tryptic peptides were found to match the cloned protein sequence. Transient expression of the cloned cDNA linked with a myc tag in COS-7 cells demonstrated high SMase activity, with an optimal pH at 9.0 and a specific dependence on taurocholate and taurochenodeoxycholate. The expressed protein reacted with both anti-myc and anti-human alk-SMase antibodies. Northern blotting of rat tissues revealed high levels of mRNA in jejunum but not in other tissues. In conclusion, we cloned rat alk-SMase cDNA from rat intestine, adjusted the putative rat alk-SMase protein in GenBank, and confirmed the specific expression of the gene in the small intestine.

Amino Acid Sequence↗

Computer generation and statistical analysis of a data bank of protein sequences translated from GenBank.

We describe PGtrans, a new and freely available protein sequence databank (2625 sequences, 554198 amino-acids). This data bank is routinely produced by automatic computer translation of the nucleotide sequence library GenBank. The information needed for the translation process (transcriptional orientation, location of coding regions, splice sites and pertinent genetic code) is gathered by the translation program through an "intelligent" scanning of the documentary field of each GenBank entry. Inconsistencies resulting in unexpected termination codons are detected and reported thus allowing the correction of data bank errors. PGtrans is intended as a tool for protein similarity searches. Its reasonable overall size (2 Moctets) makes it suitable for micro-computer environments. Up to date amino-acid composition data and relative abundances of di-, tri-, and tetra-peptides in proteins of known sequences are presented and discussed.

Amino Acid Sequence↗

Evaluation of RIDOM, MicroSeq, and Genbank services in the molecular identification of Nocardia species.

The molecular identification of Nocardia species, when compared to phenotypic identification, has two primary advantages: rapid turn-around time and improved accuracy. The information content in the 5'-end of the 16S ribosomal RNA gene is sufficient for identification of most bacterial species. An evaluation was performed to demonstrate the quality of results provided by two specialized databases (RIDOM and MicroSeq 500 versions 1.1 and 1.4.3, library version 500-0125, respectively) and the more general GenBank database. In addition, these results were compared with phenotypic identifications. Partial 5'-16S rDNA sequences from 64 culture collection strains (DSM, CIP, JCM, and ATCC) were derived, in duplicate, independently in two laboratories. Furthermore, the sequences and the conventional identification results of 91 clinical Nocardia isolates were determined. With the exception of N. soli and N. cummidelens, all Nocardia type strains were distinguishable using 5'-16S rDNA sequencing. Assuming a normal distribution for the pairwise distances of all unique Nocardia sequences and choosing a reporting criterion of > or = 99.12% similarity for a "distinct species", a statistical error probability of 1.0% can be calculated. When the various databases were searched with the clinical isolate sequences RIDOM gave a perfect match in 71.4% of cases whereas MicroSeq yielded a perfect match in only 26.4%. The GenBank service gave a 100% similarity in 59.3% but in 70.4% of these cases the results obtained were not exclusive for a single Nocardia species. Conventional methods gave a correct identification in 59 cases, although most recent taxonomic changes were not taken into account. The RIDOM service (http://www.ridom-rdna.de/) is in the process of making available a comprehensive and high-quality database for bacterial identification purposes and provides excellent results for the majority of Nocardia isolates.

DNA, Bacterial↗

Promoter Extraction from GenBank (PEG): automatic extraction of eukaryotic promoter sequences in large sets of genes.

UNLABELLED: Promoter Extraction from GenBank (PEG) extracts promoter sequences for large sets of genes using information present in GenBank. For a gene whose promoter sequence is not found, PEG will attempt to extract promoter sequences of the orthologous genes instead. AVAILABILITY: It is freely available to academic users at ftp://cshl.org/pub/science/mzhanglab/theresa/. CONTACT: zhangt@cshl.org; mzhang@cshl.org

Databases, Nucleic Acid↗

The GenBank genetic sequence databank.

The GenBank Genetic Sequence Data Bank contains over 5700 entries for DNA and RNA sequences that have been reported since 1967. This paper briefly describes the contents of the database, the forms in which the database is distributed, and the services we offer to scientists who use the GenBank database.

Animals↗

The GenBank genetic sequence data bank.

The GenBank Genetic Sequence Data Bank contains nearly 15,000 entries for DNA and RNA sequences that have been reported since 1967. This paper briefly describes the contents of the database, the forms in which the data are distributed, and the services available to scientists using the GenBank database.

Base Sequence↗

Sequence errors described in GenBank: a means to determine the accuracy of DNA sequence interpretation.

The accuracy of nucleic acid sequence data interpretation was determined by assessing and quantifying the discrepancies reported in the GenBank database. This permitted the calculation of an Error Rate (ER) for nucleic acid sequence determination. If one assumes that most entries (TB, Total Bases) were independently verified or those without reported discrepancies were correct, the ER is 0.368 errors per 1000 bases. However, if one assumes that only those sequences with reported discrepancies (TBIQ, Total Bases from entries In Question) were verified and are thus correct, the ER is 2.887 errors per 1000 bases. This establishes the first set of limit boundaries of the ER for sequence interpretation and sequence errors within the GenBank database and provides the foundation for future assessments and the monitoring of sequence data accumulation. In addition, the ER measure provides a basis to evaluate the efficiency and merit of present and future automated nucleic acid sequencing technologies which will have a direct impact upon the final outcome of the "Human Genome Initiative".

Base Sequence↗

Recent changes in the GenBank On-line Service.

The GenBank On-line Service provides access to the GenBank and EMBL nucleic acid sequence databases and to the Swiss-Prot and GenPept protein sequence databases. Users can query the databases by sequence similarity and annotation keywords and retrieve entries of interest. This access is available through e-mail servers, anonymous FTP, anonymous interactive login, and login to established, password-protected, individual accounts.

Gene Library↗

GenBank.

The GenBank nucleotide sequence database now contains sequence data and associated annotation corresponding to 56,000,000 nucleotides in 45,000 entries. The input stream of data coming into the database has largely been shifted to direct submissions from the scientific community on electronic media. The data have been installed in a relational database management system and are made available in this form through on-line access, and through various network and off-line computer-readable media. In addition, GenBank provides the U.S. distribution center for the BIOSCI electronic bulletin board service.

Base Sequence↗

GenBank.

The GenBank nucleotide sequence database now contains sequence data and associated annotation corresponding to 85,000,000 nucleotides in 67,000 entries from a total of 3,000 organisms. The input stream of data coming into the database is primarily as direct submissions from the scientific community on electronic media, with little or no data being keyboarded from the printed page by the databank staff. The data are maintained in a relational database management system and are made available in flatfile form through on-line access, and through various network and off-line computer-readable media. The data are also distributed in relational form through satellite copies at a number of institutions in the U.S. and elsewhere. In addition, GenBank provides the U.S. distribution center for the BIOSCI electronic bulletin board service.

Animals↗

GenBank.

The GenBank sequence database has undergone an expansion in data coverage, annotation content and the development of new services for the scientific community. In addition to nucleotide sequences, data from the major protein sequence and structural databases, and from U.S. and European patents is now included in an integrated system. MEDLINE abstracts from published articles describing the sequences provide an important new source of biological annotation for sequence entries. In addition to the continued support of existing services, new CD-ROM and network-based systems have been implemented for literature retrieval and sequence similarity searching. Major releases of GenBank are now more frequent and the data are distributed in several new forms for both end users and software developers.

Amino Acid Sequence↗

GenBank.

The GenBank sequence database incorporates DNA sequences from all available public sources, primarily through the direct submission of sequence data from authors and from large-scale sequencing projects. Data exchange with the EMBL Data Library and the DNA Data Bank of Japan helps ensure comprehensive coverage. GenBank continues to focus on quality control and annotation while expanding data coverage and retrieval services. An integrated retrieval system, known asEntrez, incorporates data from the major DNA and protein sequence databases, along with genome maps and protein structure information. MEDLINE abstracts from published articles describing the sequences are also included as an additional source of biological annotation. Sequence similarity searching is offered through the BLAST family of programs. All of NCBI's services are offered through the World Wide Web. In addition, there are specialized server/client versions as well as FTP and e-mail server access.

Amino Acid Sequence↗

GenBank.

The GenBank(R) sequence database (http://www.ncbi.nlm.nih.gov/) incorporates DNA sequences from all available public sources, primarily through the direct submission of sequence data from individual laboratories and from large-scale sequencing projects. Most submitters use the BankIt (WWW) or Sequin programs to send their sequence data. Data exchange with the EMBL Data Library and the DNA Data Bank of Japan helps ensure comprehensive worldwide coverage. GenBank data is accessible through NCBI's integrated retrieval system, Entrez , which integrates data from the major DNA and protein sequence databases along with taxonomy, genome and protein structure information. MEDLINE(R) abstracts from published articles describing the sequences are also included as an additional source of biological annotation. Sequence similarity searching is offered through the BLAST series of database search programs. In addition to FTP, e-mail and server/client versions of Entrez and BLAST, NCBI offers a wide range of World Wide Web retrieval and analysis services of interest to biologists.

Animals↗

Distribution of hammerhead and hammerhead-like RNA motifs through the GenBank.

Hammerhead ribozymes previously were found in satellite RNAs from plant viroids and in repetitive DNA from certain species of newts and schistosomes. To determine if this catalytic RNA motif has a wider distribution, we decided to scrutinize the GenBank database for RNAs that contain hammerhead or hammerhead-like motifs. The search shows a widespread distribution of this kind of RNA motif in different sequences suggesting that they might have a more general role in RNA biology. The frequency of the hammerhead motif is half of that expected from a random distribution, but this fact comes from the low CpG representation in vertebrate sequences and the bias of the GenBank for those sequences. Intriguing motifs include those found in several families of repetitive sequences, in the satellite RNA from the carrot red leaf luteovirus, in plant viruses like the spinach latent virus and the elm mottle virus, in animal viruses like the hepatitis E virus and the caprine encephalitis virus, and in mRNAs such as those coding for cytochrome P450 oxidoreductase in the rat and the hamster.

Amino Acid Motifs↗