Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “GenBank”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Differential expression of endogenous ferritin genes and iron homeostasis alteration in transgenic tobacco overexpressing soybean ferritin gene.

For studying the effects of endogenous ferritin gene expressions (NtFer1, GenBank accession number AY083924; and NtFer2, GenBank accession number AY141105) on the iron homeostasis in transgenic tobacco (Nicotiana tabacum L.) plants expressing soybean (Glycine max Merr) ferritin gene (SoyFer1, GenBank accession number M64337), the transgenic tobacco has been produced by placing soybean ferritin cDNA cassette under the control of the CaMV 35S promoter. The exogenous gene expression was examined by both Northern- and Western-blot analyses. Comparison of endogenous ferritin gene expressions between nontransformant and transgenic tobacco plants showed that the expression of NtFer1 was increased in the leaves of transgenic tobacco plants, whereas the NtFer2 expression was unchanged. The iron concentration in the leaves of transgenic tobacco plants was about 1.5-folds higher than that in nontransformant. Enhanced growth of transgenic tobacco was observed at the early development stages, resulting in plant height and fresh weights significantly greater than those in the nontransformant. These results demonstrated that exogenous ferritin expression induced increased expression of at least one of the endogenous ferritin genes in transgenic tobacco plants by enhancing the ferric chelate reductase activity and iron transport ability of the root, and improved the rate of photosynthesis.

FMN Reductase↗

[In silico cloning of Efp-0, a novel earthworm fibrinolytic enzyme gene and verification of its coding region by RT-PCR].

There are four different types of N-terminal amino acid sequences (F-I-0, F-I, F-II, F-III) in the multicomponents of earthworm fibrinolytic enzymes (EFE). In GenBank 21 nucleic acid sequences of EFE have been reported. Among them, most of the N-terminal amino acid sequences belong to the F-III type,few belong to the F-II type. Only one is similar to the F-I type, but none to F-I-0. In this research we hoped to obtain the gene encoding component F-I-0 of EFE by the bioinformatics tools. Based on the N-terminal amino acid sequence VVGGSDTTIGQYPHQL of the F-I-0 type from Lumbricus rubellus, a nucleic acid sequence was obtained by in silico cloning from dbEST of Lumbricidae using the software DNAMAN. A new gene of EFE from Eisenia foetida was successfully obtained by RT-PCR using specific primers designed according to this sequence. The new gene named EfP-0 was cloned in pMAL-c2x and expressed as the fusion protein MBP-EfP-0 in the supernatant of lysate. The fusion protein MBP-EfP-0 purified by affinity chromatography had hydrolytic activity on casein plate. Sequencing result shows, EfP-0 has 678bp and encodes a protein of 225 amino acids. The protein is a serine protease belonging to trypsin family. It has similar amino acid composition to F-I-0. BLAST in GenBank shows that the similarity is lower than 40% between EJP-0 gene and other EFE genes. By this we conclude that EfP-0 gene of EFE is a novel gene and it is the first time to be reported, its accession number for Genbank is DQ836917.

Amino Acid Sequence↗

A mRNA molecule encoding truncated excitatory amino acid carrier 1 (EAAC1) protein (EAAC2) is transcribed from an independent promoter but not an alternative splicing event.

Glutamate transporter EAAC1 removes excitatory neurotransmitter in central nervous system, and also absorbs glutamate in epithelia of intestine, kidney, liver and heart for normal cell growth. When a mouse cDNA was screened using EAAC1 cDNA fragment as probe in our lab, a transcript (GenBank U75214) encoding an EAAC1 protein with 148 residues truncated at N-terminal was cloned and named as EAAC2. Sequence analysis shows that EAAC2 has it's own start code and unique 5'UTR that is different from that of EAAC1. A mouse genomic library was screened and a positive clone including EAAC1 CDS was sequenced (GenBank AF 322393) and indicates that normal EAAC1 transcript (GenBank U73521) is transcribed from 10 exons in terms of exon I, II, III, IV, V, VI, VII, VIII, IX, X, and EAAC2 transcript is consisted by exons from IV to IX as same as that of EAAC1 and with its unique exon beta upstream to exon IV and exon delta downstream to IX. EAAC2 transcript has a cluster of transcriptional start sites not overlapping with the transcriptional start sites of EAAC1. These results indicate that EAAC2 is transcribed from an independent promoter but not an alternative splicing event.

5' Untranslated Regions↗

Two polymorphisms within interleukin-3 (hIL3) gene detected by mismatch PCR/RFLP.

Two alleles of IL-3 have been reported to GenBank (GenBank M14743, M20137). The sequence difference between these two alleles is at the first nucleotide of the 27th codon (the 131st nucleotide from the initiation site): thymine and cytosine, and leading the amino acid difference: proline and serine (Pro27Ser). The other allelism, thymine and cytosine, was also observed at position -16 of the IL-3 upstream promotor region (GenBank L10616, M60870). We clarified that these substitutions were frequent polymorphisms in the Japanese population by using the mismatch-PCR (polymerase chain reaction)/RFLP (restriction fragment length polymorphism) method.

Alleles↗

Comparative sequencing and association studies of aromatic L-amino acid decarboxylase in schizophrenia and bipolar disorder.

Aromatic L-amino acid decarboxylase (AADC) is a relatively non specific enzyme involved in the biosynthesis of several classical neurotransmitters including dopamine and 5-hydroxytryptamine (5HT; serotonin). AADC does not catalyse the rate limiting step in either pathway, but is rate limiting in the synthesis of 2-phenylethylamine (2PE) which is a positive modulator of dopaminergic transmission and a candidate natural psychotogenic compound.1 We and others have proposed that polymorphism in AADC resulting in altered 2PE activity might contribute to the pathogenesis of psychosis. In order to test this hypothesis, we have used denaturing high performance liquid chromatography (DHPLC)3 to screen 3943 bases of the AADC gene and its promoter regions for variants that might affect protein structure or expression in 15 unrelated people with schizophrenia, and 15 unrelated people with bipolar disorder. Three polymorphisms were identified by DHPLC: a insertion/deletion polymorphism in the 5' UTR of the neuronal specific mRNA (g.-33-30delAGAG, bases 586-589 of GenBank M77828), a T>A variant in the non-neuronal exon 1 (g. -67T>A, GenBank M88070), and a G>A polymorphism within intron 8 (g. IVS8 +75G>A, GenBank M84598). Case-control analysis did not suggest that genetic polymorphism in the AADC gene is associated with liability for developing schizophrenia or bipolar disorder.

Aromatic-L-Amino-Acid Decarboxylases↗

Molecular cloning, bacterial expression and properties of Rab31 and Rab32.

GTP-binding proteins of the Rab family were cloned from human platelets using RT-PCR. Clones corresponding to two novel Rab proteins, Rab31 and Rab32, and to Rab11A, which had not been detected in platelets previously, were isolated. The coding sequence of Rab31 (GenBank accession no. U59877) corresponded to a 194 amino-acid protein of 21.6 kDa. The Rab32 sequence was extended to 1000 nucleotides including 630 nucleotides of coding sequence (GenBank accession no. U59878) but the 5' coding sequence was only completed later by others (GenBank accession no. U71127). Human Rab32 cDNA encodes a 225 amino-acid protein of 25.0 kDa with the unusual GTP-binding sequence DIAGQE in place of DTAGQE. Northern blots for Rab31 and Rab32 identified 4.4 kb and 1.35 kb mRNA species, respectively, in some human tissues and in human erythroleukemia (HEL) cells. Rabbit polyclonal anti-peptide antibodies to Rab31, Rab32 and Rab11A detected platelet proteins of 22 kDa, 28 kDa and 26 kDa, respectively. Human platelets were highly enriched in Rab11A (0.85 microg x mg of platelet protein(-1)) and contained substantial amounts of Rab32 (0.11 microg x mg protein(-1)). Little Rab31 was present (0.005 microg x mg protein(-1)). All three Rab proteins were found in both granule and membrane fractions from platelets. In rat platelets, the 28-kDa Rab32 was replaced by a 52-kDa immunoreactive protein. Rab31 and Rab32, expressed as glutathione S-transferase (GST)-fusion proteins, did not bind [alpha-(32)P]GTP on nitrocellulose blots but did bind [(35)S]GTP[S] in a Mg(2+)-dependent manner. Binding of [(35)S]GTP[S] was optimal with 5 microm Mg(2+)(free) and was markedly inhibited by higher Mg(2+) concentrations in the case of GST-Rab31 but not GST-Rab32. Both proteins displayed low steady-state GTPase activities, which were not inhibited by mutations (Rab31(Q64L) and Rab32(Q85L)) that abolish the GTPase activities of most low-M(r) GTP-binding proteins.

Amino Acid Sequence↗

Mapping of expressed sequence tags from a porcine early embryonic cDNA library.

The goal of this study was to identify and map genes expressed during the elongation phase of embryogenesis in swine. Expressed sequence tags were analysed from a previously described porcine cDNA library prepared from elongating swine embryos. Average insert length of randomly selected clones was approximately 600 bp, with a range from < 100 to > 2500 bp. Single-pass, coding strand sequences from 1132 independent clones were compared with the GenBank non-redundant (nr) database via BLASTN analysis to identify potential porcine homologous of known genes. Among these sequences, 781 (69%) showed significant (score > 300) homology to non- mitochondrial sequences previously deposited in GenBank. Sequences matching interleucin 1 beta and thymosin beta 10 were most frequently observed (24 and 18 clones, respectively), in addition to matches with 310 other distinct genes. No significant match in the GenBank nr database was obtained for 303 sequences. Analysis demonstrated that 151 (50%) had open reading frames (ORF) extending at least 50 codons from the first base of the clone insert. Genetic markers were developed and used to map a subset of 17 genes, selected on the basis of function or of the ability to design primers that successfully amplified porcine genomic DNA, to 10 different porcine chromosomes, providing a set of mapped markers corresponding to genes expressed during conceptus elongation.

Alleles↗

Putative in silico mapping of DNA sequences to livestock genome maps using SSLP flanking sequences.

In this study, an in silico approach was developed to identify homologies existing between livestock microsatellite flanking sequences and GenBank nucleotide sequences. Initially, 1955 bovine, 1570 porcine and 1121 chicken microsatellites were downloaded and the flanking sequences were compared with the nr and dbEST databases of GenBank. A total of 74 bovine, 44 porcine and 37 chicken microsatellite flanking sequences passed our criteria and had at least one significant match to human genomic sequence, genes/expressed sequence tags (ESTs) or both. GenBank annotation and BLAT searches of the UCSC human genome assembly revealed that 38 bovine, 13 porcine and 17 chicken microsatellite flanking sequences were highly similar to known human genes. Map locations were available for 67 bovine, 44 porcine and 21 chicken microsatellite flanking sequences, providing useful links in the comparative maps of humans and livestock. In support of our approach, 112 alignments with both microsatellite and match mapping information were located in the expected chromosomal regions based on previously reported syntenic relationships. The development of this in silico mapping approach has significantly increased the number of genes and EST sequences anchored to the bovine, porcine and chicken genome maps and the number of links in various human-livestock comparative maps.

Animals↗

Two bi-allelic single nucleotide polymorphisms within the promoter region of the horse tumour necrosis factor alpha gene.

Primers based on GenBank sequences within the 5' untranslated region (UTR) of the human and horse tumour necrosis factor alpha (TNF-alpha) genes were designed and used to amplify a 522-bp product. Sequencing of five clones derived from five independent PCRs obtained from three different animals of three different breeds (Old Kladruber, Akhal-Teke and Shetland Pony) revealed a high level of sequence identity to the TNF-alpha promoter regions of other species. The existing GenBank horse sequences were confirmed and extended upstream by 230 nucleotides. Based on the sequence obtained, a new horse-specific forward primer was designed to amplify a 213-bp PCR product, which was screened for polymorphism using single-strand conformation polymorphism (SSCP). Three allelic variants of the horse TNF-alpha gene were identified and sequenced (GenBank accession numbers ADF 349558-60). Two single nucleotide polymorphisms explained the existence of the three SSCP alleles detected: C/T and T/C single base pair substitutions at positions 137 and 147, respectively. Differences in allelic frequencies between Old Kladruber and Akhal-Teke breeds were observed.

5' Untranslated Regions↗

Expression of functional Anopheles merus alpha-amylase in the baculovirus/Spodoptera frugiperda system.

The Anopheles merus (Diptera, Nematocera, Culicoidea) alpha-amylase gene (AmerAmy, GenBank Accession Number U01210) was amplified with its own or with the Zabrotes subfasciatusalpha-amylase signal peptide (ZsAmerAmy, GenBank Accession Number AY270183) by PCR, using designed primers. The AmerAmy gene was sequenced from its promotor to the TGA codon. As a positive control, the Z. subfasciatusalpha-amylase gene with its own signal peptide (ZsAmy, GenBank Accession Number AF255722) was also amplified by PCR. These three sequences were inserted into the baculovirus genome using the Bac-to-Bac trade mark system. Recombinant baculovirus preparations were used to infect Sf9 Spodoptera frugiperda insect cells. The A. merusalpha-amylase was successfully expressed as an active enzyme detected mainly in cell culture supernatants.

Animals↗

Coding properties of macronuclear DNA molecules in Sterkiella nova (Oxytricha nova).

The DNA in the macronucleus of the stichotrichs like Sterkiella nova (formerly Oxytricha nova) occurs in short molecules ranging from approximately 200 bp to approximately 20,000 bp. It has been estimated that there are approximately 24,500 different sized DNA molecules in the macronucleus. Single genes have been assigned to approximately 130 different sized macronuclear molecules in various stichotrichs (12 in Sterkiella nova) and hypotrichs, suggesting that each of the -24,500 different sized molecules encodes a different gene. To test this proposition we sequenced 31 macronuclear molecules picked randomly from a plasmid library of macronuclear DNA and analyzed them for potential gene content. The open reading frames (ORFs) in three short molecules encode amino acid (aa) sequences that do not match sequences in GenBank. They may or may not encode genes. Twenty-eight of the 31 molecules contain ORFs encoding aa sequences with significant matches to sequences in GenBank. Six molecules contain more than one ORF with a significant match to GenBank. These results indicate that almost all, if not all of the -24,500 different molecules encode one or more genes, yielding an estimate of -26,800 genes in the macronucleus of S. nova.

Animals↗

Molecular and morphological characterization of Aedes albopictus in northwestern Greece and differentiation from Aedes cretinus and Aedes aegypti.

The presence of Aedes albopictus (Skuse) was recently confirmed for the first time in northwestern Greece. This location is within the distribution range of a morphologically similar species, Aedes cretinus Edwards, and is a potentially favorable region for the reintroduction of Aedes aegypti (L.). It was thus compelling to use methods in addition to morphology-based keys to correctly identify specimens badly damaged, rubbed, or otherwise altered in their external characteristics. It was decided to use molecular techniques as a novel and reliable method for differentiating the three Stegomyia species. The nuclear internal transcribed spacer 2 (ITS2) fragments from morphologically identified Ae. albopictus and Ae. cretinus specimens were amplified, and their sequences were compared with those in GenBank for Ae. albopictus, Ae. cretinus, and Ae. aegypti. Also, mitochondrial cytochrome oxidase I (COI) fragments were amplified for Ae. albopictus and Ae. cretinus (so far not available in GenBank) and compared with Ae. aegypti fragments. ITS2 and COI sequences generated in our study were deposited in GenBank and could be useful in future studies of mosquitoes by other research workers.

Aedes↗

GENEVIEW and the DNACE data bus: computational tools for analysis, display and exchange of genetic information.

We describe an interactive computational tool, GENEVIEW, that allows the scientist to retrieve, analyze, display and exchange genetic information. The scientist may request a display of information from a GenBank locus, request that a restriction map be computed, stored and superimposed on GenBank information, and interactively view this information. GENEVIEW provides an interface between the GenBank data base and the programs of the Lilly DNA Computing Environment (DNACE). This interface stores genetic information in a simple, free format that has become the universal convention of DNACE; this format will serve as the convention for all future software development at Eli Lilly and Company, and could serve as a convention for genetic information exchange.

Animals↗

Corruption of genomic databases with anomalous sequence.

We describe evidence that DNA sequences from vectors used for cloning and sequencing have been incorporated accidentally into eukaryotic entries in the GenBank database. These incorporations were not restricted to one type of vector or to a single mechanism. Many minor instances may have been the result of simple editing errors, but some entries contained large blocks of vector sequence that had been incorporated by contamination or other accidents during cloning. Some cases involved unusual rearrangements and areas of vector distant from the normal insertion sites. Matches to vector were found in 0.23% of 20,000 sequences analyzed in GenBank Release 63. Although the possibility of anomalous sequence incorporation has been recognized since the inception of GenBank and should be easy to avoid, recent evidence suggests that this problem is increasing more quickly than the database itself. The presence of anomalous sequence may have serious consequences for the interpretation and use of database entries, and will have an impact on issues of database management. The incorporated vector fragments described here may also be useful for a crude estimate of the fidelity of sequence information in the database. In alignments with well-defined ends, the matching sequences showed 96.8% identity to vector; when poorer matches with arbitrary limits were included, the aggregate identity to vector sequence was 94.8%.

Base Sequence↗

Codon usage tabulated from the international DNA sequence databases.

CUTG (codon usage tabulated from GenBank) is a comprehensive database for codon usage. The codon usage for each full-length protein gene has been calculated using the nucleotide sequence obtained from GenBank sequence database. The sum of the codon use of each organism has been also calculated. The data files can be obtained from anonymous ftp sites of DDBJ, DISC and EBI. The list of codonusage of genes in organisms was made searchableby name of organism through a web site http://www.dna.affrc.go.jp/ approximately nakamura/CUTG.html The compilation is synchronized with major release of GenBank.

Animals↗

DAtA: database of Arabidopsis thaliana annotation.

The Database of Arabidopsis thaliana Annotation (D At A) was created to enable easy access to and analysis of all the Arabidopsis genome project annotation. The database was constructed using the completed A.thaliana genomic sequence data currently in GenBank. An automated annotation process was used to predict coding sequences for GenBank records that do not include annotation. D At A also contains protein motifs and protein similarities derived from searches of the proteins in D At A with motif databases and the non-redundant protein database. The database is routinely updated to include new GenBank submissions for Arabidopsis genomic sequences and new Blast and protein motif search results. A web interface to D At A allows coding sequences to be searched by name, comment, blast similarity or motif field. In addition, browse options present lists of either all the protein names or identified motifs present in the sequenced A.thaliana genome. The database can be accessed at http://baggage. stanford.edu/group/arabprotein/

Arabidopsis↗

Analysis of canonical and non-canonical splice sites in mammalian genomes.

A set of 43 337 splice junction pairs was extracted from mammalian GenBank annotated genes. Expressed sequence tag (EST) sequences support 22 489 of them. Of these, 98.71% contain canonical dinucleotides GT and AG for donor and acceptor sites, respectively; 0.56% hold non-canonical GC-AG splice site pairs; and the remaining 0.73% occurs in a lot of small groups (with a maximum size of 0.05%). Studying these groups we observe that many of them contain splicing dinucleotides shifted from the annotated splice junction by one position. After close examination of such cases we present a new classification consisting of only eight observed types of splice site pairs (out of 256 a priori possible combinations). EST alignments allow us to verify the exonic part of the splice sites, but many non-canonical cases may be due to intron sequencing errors. This idea is given substantial support when we compare the sequences of human genes having non-canonical splice sites deposited in GenBank by high throughput genome sequencing projects (HTG). A high proportion (156 out of 171) of the human non-canonical and EST-supported splice site sequences had a clear match in the human HTG. They can be classified after corrections as: 79 GC-AG pairs (of which one was an error that corrected to GC-AG), 61 errors that were corrected to GT-AG canonical pairs, six AT-AC pairs (of which two were errors that corrected to AT-AC), one case was produced from non-existent intron, seven cases were found in HTG that were deposited to GenBank and finally there were only two cases left of supported non-canonical splice sites. If we assume that approximately the same situation is true for the whole set of annotated mammalian non-canonical splice sites, then the 99.24% of splice site pairs should be GT-AG, 0.69% GC-AG, 0.05% AT-AC and finally only 0.02% could consist of other types of non-canonical splice sites. We analyze several characteristics of EST-verified splice sites and build weight matrices for the major groups, which can be incorporated into gene prediction programs. We also present a set of EST-verified canonical splice sites larger by two orders of magnitude than the current one (22 199 entries versus approximately 600) and finally, a set of 290 EST-supported non-canonical splice sites. Both sets should be significant for future investigations of the splicing mechanism.

Animals↗

Association of polymorphisms for prolactin and prolactin receptor genes with broody traits in chickens.

Prolactin (PRL) is generally accepted as crucial to the onset and maintenance of broodiness in avian species. The prolactin receptor (PRLR) plays an important role in the PRL signal transduction cascade. Two candidate genes, PRL and PRLR, were screened for polymorphisms in the chicken, and their genetic effects on broodiness were evaluated. Pedigreed hens (n = 155) of the Blue-shell chicken, a Chinese local breed, were observed for phenotypic broody traits including nesting days, broody days, repeats of broody cycles, and duration of broodiness. For polymorphism analysis, White Leghorns, Hy-Line brown egg layers, Avian broilers, and some other Chinese local breeds were included. Fifteen sets of primers were used to amplify the nucleotide sequences of the promotor of PRL and exons of PRLR. The PCR products were screened for polymorphisms using single-stranded conformational polymorphism protocol. Sequencing revealed a 24-bp insertion occurring in the promotor, -377 approximately -354, of PRL (GenBank accession no. AB011434). A single nucleotide polymorphism (SNP), A9026G (GenBank accession no. AY237377), in exon 3 of PRLR was also detected, which led to a nucleotide transition in the 5'-untranslated region (5'-UTR) of PRLR cDNA. Two SNP, T14771C and G14820A (GenBank accession no. AY237376), were detected in exon 6 of the PRLR. The T14771C transition led to an amino acid variation, Leu340Ser, in PRLR, whereas the G14820A transition was a synonymous mutation. An association analysis showed that the genetic polymorphisms at PRLR3 and PRLR6 were not related to broodiness (P > 0.05), whereas the individuals without the insertion sequence at PRLpro2 were associated with broody traits (P < 0.05) and the incidence (>30%) of typical broody of genotypes +/- and -/- was higher (P < 0.01) than that of +/+. In addition, all White Leghorns were +/+ for PRLpro2, whereas local breeds with very strong broodiness were nearly all -/-. Homozygous insertion of the 24-bp sequence in the PRL promoter may decrease the expression of PRL, leading to nonbroodiness. The results suggested that PRLpro2 could be a genetic marker in breeding against broodiness in chickens.

Animals↗