Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “GenBank”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Isolation of the human BACH1 transcription regulator gene, which maps to chromosome 21q22.1.

In order to contribute to the development of the transcriptional map of chromosome 21, we performed exon trapping using cosmid clones mapped in the region 21q22.1-22.2 and identified a number of potential exons. One of the trapped exons (Genbank No. AF026200) showed a strong homology with the mouse Bach1 gene (Genbank No. D86603), a transcription factor regulating gene expression. We then isolated the full-length coding region of the human BACH1 gene using expressed sequence tags, reverse transcription-polymerase chain reaction and rapid amplification of cDNA ends. The predicted BACH1 protein contains 736 amino acids and is 88% identical to its mouse homolog. It contains basic leucine zipper and BTB-zinc finger domains (which are directly involved in DNA binding for transcription regulation). The BACH1 gene maps in a relatively gene-poor region on 21q22.1 in yeast artificial chromosome 814c1 of the collection of Chumakov et al. Northern blot analysis revealed that it is expressed as an mRNA species of approximately 5.8 kb in all 16 adult and 4 fetal tissues examined; an additional mRNA species of 2.8 kb was observed in adult testis. The contribution of the BACH1 gene to the pathophysiology of trisomy or monosomy 21 is unknown. In addition, no monogenic disorders associated with mutations in the BACH1 gene have yet been identified.

Adult↗

Yellow leaf of sugarcane is caused by at least three different genotypes of sugarcane yellow leaf virus, one of which predominates on the Island of Réunion.

The genetic diversity of sugarcane yellow leaf virus (SCYLV) was analyzed with 43 virus isolates from Réunion Island and 17 isolates from world-wide locations. We attempted to amplify by reverse-transcription polymerase chain reaction (RT-PCR), clone, and sequence four different fragments covering 72% of the genome of these virus isolates. The number of amplified isolates and useful sequence information varied according to each fragment, whereas an amplicon was obtained with diagnostic primers for 59 out of 60 isolates (98%). Phylogenetic analyses of the sequences determined here and additional sequences of 11 other SCYLV isolates available from GenBank showed that SCYLV isolates were distributed in different phylogenetic groups or belonged to single genotypes. The majority of isolates from Réunion Island were grouped in phylogenetic clusters that did not contain any isolates from other origins. The complete six ORFs (5612 bp) of five SCYLV isolates (two from Réunion Island, one from Brazil, one from China, and one from Peru) were amplified, cloned, and sequenced. The existence of at least three distinct genotypes of SCYLV was shown by phylogenetic analysis of the sequences of these isolates and additional published sequences of three SCYLV isolates (GenBank accessions). The biological significance of these genotypes and of the origin of the distinct lineage of SCYLV in Réunion Island remains to be determined.

Cloning, Molecular↗

Novel interferon regulatory factor-1 polymorphisms in a Kenyan population revealed by complete gene sequencing.

Variation in susceptibility to HIV-1 infection depends on numerous factors, and host genetic variation has been well-described as an important component. As a transcriptional regulator, interferon regulatory factor 1 (IRF-1) plays a key role in both innate and adaptive immunity against viral infection. IRF-1 has also been shown to directly interact with HIV-1 5' LTR and efficiently initiate or amplify HIV-1 replication. By complete gene sequencing, we investigated genetic polymorphism of the IRF-1 gene in an HIV-1-endemic Kenyan population. This population displayed extensive genetic diversity at the IRF-1 locus. Fifty-three single nucleotide polymorphisms (SNPs) were identified in this population, including 26 novel SNPs. Two insertion and one deletion polymorphisms in IRF-1 were also identified. Linkage disequilibrium (LD) among these genetic variations was shown to be common in IRF-1. The functional consequences of these mutations in the context of HIV-1/AIDS remain to be determined. We also identified 35 consistent discrepancies between IRF-1 GenBank sequences and our population based sequencing data, suggesting that the previously submitted GenBank data were not representative of the majority of human IRF-1 sequences.

Base Sequence↗

Identification of immune-related genes in hemocytes of black tiger shrimp (Penaeus monodon).

An expressed sequence tag (EST) library was constructed from hemocytes of the black tiger shrimp (Penaeus monodon) to identify genes associated with immunity in this economically important species. The number of complementary DNA clones in the constructed library was approximately 4 x 10(5). Of these, 615 clones having inserts larger than 500 bp were unidirectionally sequenced and analyzed by homology searches against data in GenBank. Significant homology to known genes was found in 314 (51%) of the 615 clones, but the remaining 301 sequences (49%) did not match any sequence in GenBank. Approximately 35% of the matched ESTs were significantly identified by the BLASTN and BLASTX programs, while 65% were recognized only by the BLASTX program. Of the 615 clones, 55 (8.9%) were identified as putative immune-related genes. The isolated genes were composed of those coding for enzymes and proteins in the clotting system and the prophenoloxidase-activating system, antioxidative enzymes, antimicrobial peptides, and serine proteinase inhibitors. Three full-length ESTs encoding antimicrobial peptides (antilipopolysaccharide and penaeidin homologues) and a heat shock protein (cpn10 homologue) are reported.

Journal Article↗

Analysis of expressed sequence tags from calcifying cells of marine coccolithophorid (Emiliania huxleyi).

An expressed sequence tag (EST) approach was used to investigate gene expression in the unicelluar marine alga Emiliania huxleyi. We randomly selected 3000 EST sequences from a cDNA library of transcripts expressed under conditions promoting coccolithogenesis. Cluster analysis and contig assembly resulted in a unigene set of approximately 1523 ESTs. Only 36% of the unique sequences exhibited significant homology to sequences in GenBank. Of particular interest were the numerous transcripts with homology to sequences associated with sexual reproduction and calcium homeostasis in other unicellular and multicellular organisms. The majority of ESTs (64%) had little or no significant sequence homology to entries in GenBank, suggesting a potential for further novel gene discovery. The catalog of ESTs reported herein represents a significant increase in the limited sequence information currently available for E. huxleyi and should make the coccolithophorid more accessible to powerful genomics and postgenomics technologies.

Base Composition↗

Development of expressed sequence tags from the bay scallop, Argopecten irradians irradians.

The bay scallop, Argopecten irradians irradians, introduced from North America, has become one of the most important aquaculture species in China. Inan effort to identify scallop genes involved in host defense, a high-quality cDNA library was constructed from whole body tissues of the bay scallop. A total of 5828 successful sequencing reactions yielded 4995 expressed sequence tags (ESTs) longer than 100 bp. Cluster and assembly analyses of the ESTs identified 637 contigs (consisting of 2853 sequences) and 2142 singletons, totaling 2779 unique sequences. Basic Local Alignment Search Tool (BLAST) analysis showed that the majority (73%) of the unique sequences had no significant homology (E-value </= 0.005) to sequences in GenBank. Among the 748 sequences with significant GenBank matches, 160 (21.4%) were for genes related to metabolism, 131 (17.5%) for cell/organism defense, 124 (16.6%) for gene/protein expression, 83 (11.1%) for cell structure/motility, 70 (9.4%) for cell signaling/communication, 17 (2.3%) for cell division, and 163 (21.8%) matched to genes of unknown functions. The list of host-defense genes included many genes with known and important roles in innate defense such as lectins, defensins, proteases, protease inhibitors, heat shock proteins, antioxidants, and Toll-like receptors. The study provides a significant number of ESTs for gene discovery and candidate genes for studying host defense in scallops and other molluscs.

Animals↗

Molecular identification, polymorphism, and expression analysis of major histocompatibility complex class IIA and B genes of turbot (Scophthalmus maximus).

Major histocompatibility complex (MHC) class II has a central role in the adaptive immune system by presenting foreign peptides to the T-cell receptor. The full lengths of MHC class II A and B cDNA were cloned from turbot by homology cloning and rapid amplification of cDNA ends polymerase chain reaction (RACE PCR), and genomic organization, molecular polymorphism, and expression of turbot class IIB gene were examined to study the function of class IIB gene in fish. The deduced amino acid sequence of turbot class II A (GenBank accession no.DQ001730) and turbot class IIB (GenBank accession no. DQ094170) had 69.8%, 67.6%, 65.5%, 59.2%, 54.5%, 52.8%, 46.2%, 46.6%, 28.3%, 28.5%, 22.2% identity and 71.5%, 70.7%, 67.1%, 68.4%, 46.7%, 53.5%, 46.7%, 50.0%, 25.2%, 29.2%, 27.6% identity with those of Japanese flounder, striped sea bass, red sea bream, cichlid, rainbow trout, Atlantic salmon, carp, zebrafish, nurse shark, mouse and human, respectively. Eleven class IIB alleles were identified from three turbot individuals. The amino acid sequence of turbot class IIB designated as Scma-DAB*0101 had 86.9%, 88.6%, 88.6%, 89.4%, 87.8%, 86.9%, 84.1%, 86.5%, 87.3%, 77.1%, and 86.9% identity with those of turbot class IIB 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 (Scma-DAB*0201- Scma-DAB*1201), respectively. Six different class IIB alleles observed in a single individual may infer the existence of three loci at least. Semiquantitative reverse transcriptase PCR (RT-PCR) demonstrated that turbot class IIA and B were ubiquitously expressed in normal tissues. Challenge of turbot with pathogenic bacteria, Vibrio anguillarum, resulted in a significant decrease in the expression of MHC class IIB mRNA from 24 h to 48 h after infection in liver and head kidney, and a significant decrease from 24 h to 72 h after infection in spleen, followed by an increase after 96 h, respectively.

Amino Acid Sequence↗

Identification of transcriptionally regulated genes in response to cellular iron availability in rat hippocampus.

The present study was attempted to identify transcriptionally regulated genes of the normal neurocytes responsive to iron availability. Postnatal rat hippocampus cells were primarily cultured either under the iron-loaded or depleted conditions. These cultured cells were applied for the generation of subtracted complementary DNA libraries by the suppression subtraction hybridization (SSH) and for the subsequent identification of differentially expressed transcripts by reverse Northern blot. The differentially expressed genes were chosen to perform sequencing, and then some of them were performed by Northern blot analysis for observation of their expression in the hippocampus of rats with the different iron status. The results indicated that five unique transcripts were strong candidates for differential expression in cellular iron repletion, one of them is a novel sequence (GenBank No. AF 433878), while 26 unique transcripts were strong candidates for differential expression in cellular iron deprivation, one of them is a novel sequence (GenBank No. AY 912101). The revealed known genes responsive to iron availability were previously unknown to respond to iron availability, or have not been determined in the brain, have not even been currently determined in their physiological and biological functions. Interestingly, the proteins encoded by most of the known genes are either directly pointed to or indirectly associated with the molecules that play important, even key roles in cellular signal transduction and the cell cycle. These findings lead to the important suggestion that the cellular responses to iron availability involve extensive transcriptional regulation and cellular signal transduction. Therefore, iron may serve as a signal, which directly and/or indirectly regulates or modulates cell functions.

Animals↗

Revealing molecular targets for enterovirus type 71 detection by profile hidden Markov models.

The enterovirus infection in 1998 claimed 78 deaths in Taiwan, with an average of 40 fatalities each year after. Traditional serum-based diagnostic methods often fail to detect enteroviruses due to antigenic changes. As a result, many isolates remain untyped and are absent from the enterovirus surveillance and epidemiological investigations. We present a profile hidden Markov model (HMM) method for molecular typing of enterovirus 71 (EV71). Based on the enteroviral sequences retrieved from GenBank, we build a nucleotide-based and an amino acid-based profile HMM for each EV71 gene using the package HMMER. HMMER bit score-based Z-scores for EV71 and non-EV71 sequences are calculated for each of these profile HMMs. In a genome-wide analysis, we find that the distribution of the EV71 Z-scores and that of the non-EV71 Z-scores have disjoint support for nucleotide-based VP1 profile HMM if the sequence is longer than 150 bases; a VP1-based molecular typing method for EV71 is thus proposed. We also report VP4 an alternative molecular target for detecting EV71, while the two UTRs and all the genes coding the internal proteins cannot be used for such purpose. To demonstrate the performance of the nucleotide-based EV71 VP1 profile HMM, 330 enterovirus VP1 nucleotide sequences newly reported to GenBank are typed with this method. All the EV71 sequences are detected with no error.

Amino Acid Sequence↗

A novel frequent BRCA1 allele in Chinese patients with breast cancer.

The whole length of exon 11 of BRCA1 was sequenced (total 3427 bp) in 59 patients and 10 healthy female blood donors. To allow a rapid determination of the different BRCA1 alleles, a sequence-specific primer PCR method (PCR-SSP) was established and was applied to 57 additional female donors. Finally, the full-length coding region of BRCA1 was analyzed through reversed-transcriptase PCR (RT-PCR) and cDNA sequencing (total 5554 bp) in one donor with wild-type allele and 2 patients with one or two mutated alleles. By genomic DNA sequencing, 5 homozygous polymorphisms were observed in 18 patients: 2201C>T, 2430T>C, 2731C>T, 3232A>G and 3667A>G All of them were previously observed in Caucasians, Malay and Chinese, but for the first time the mutations were found in one allele (GenBank AY304547). Twenty-six patients and 4 donors were heterozygous at these 5 nucleotide positions. The remaining 15 patients and 6 donors showed a sequence identical with the standard BRCA1 gene. Combined the PCR-SSP results and in a summary, 6 of 67 (9.0 %) healthy individuals were homozygous for the mutated allele, whereas 18 of 59 (30.5 %) breast cancer patients were homozygous. A Chi-square test showed a significant correlation between homozygous mutated BRCA1 allele and breast cancer. The cDNA sequencing showed that 2 additional mutations, 4427T>C in exon 13 and 4956A>G in exon 16, were found. A new BRCA1 allele, which is BRCAI-2201T/2430C/2731T/3232G/3667G/4427C/4956G (GenBank AY751490), was found in Chinese. And the homozygote of this mutated allele may implicate a disease-association in Chinese.

Adult↗

Molecular characterization of duck hepatitis B virus isolated from Hubei brown ducks.

The objective of this study was to characterize the genome structure of duck hepatitis B virus (DHBV) isolated from Hubei brown ducks. The natural carrier rate of DHBV in adult ducks from Hubei area was investigated and the DHBV DNA-positive serum screened out. The complete genome of a DHBV strain was amplified by polymerase chain reaction (PCR) and cloned into T vector and sequenced. The results showed that the carrier rate of DHBV in Hubei brown ducks was 10 %. This strain (GenBank accession number DQ276978) had a genome of 3024 nucleotides with three overlapping open reading frames encoding the surface, core and polymerase proteins respectively. Comparison of the strain with 17 DHBV strains registered in GenBank revealed a homology from 89.3 % to 93.5 % at the nucleotide level. The sequences of the structural and functional domains of these proteins were highly conserved. The strain was found to share more signature amino acids in the polymerase genes with the "Chinese" DHBV strains than those of the "Western" country strains. This finding was also corroborated by a phylogenetic tree analysis. Therefore, the DQ276978 might belong to a subtype of the Chinese DHBV strains.

Animals↗

Global variation in G+C content along vertebrate genome DNA. Possible correlation with chromosome band structures.

The global, rather than local, variation in G+C content along the nuclear DNA sequences of various organisms was studied using GenBank sequence data. When long DNA sequences of the genomes of Escherichia coli and Saccharomyces cerevisiae were examined, the levels of their G+C content (G+C%) were found to be within a narrow range around that of the whole genome. The G+C% levels for sequences of vertebrate genomes, however, were found to cover a wide range, showing that their genome is a mosaic of sequences with different G+C% levels, in each of which the sequence is fairly homogeneous in its G+C% for a very long distance. Through surveying a human genetic map and GenBank DNA sequences, the global variations in G+C% along the human genome DNA were found to be correlated with chromosome band structures.

Animals↗

Construction of a facsimile data set for large genome sequence analysis.

A test was devised for exploring the question of whether it will be possible to identify genes in large-scale genome studies solely by sequence comparison with current sequence collections. To this end, a facsimile data set was constructed by dividing GenBank Release 56 randomly into two halves, one to serve as a reference set and the other intended to simulate raw data anticipated from large genome sequence projects. All supplementary information and identifying marks were removed from the test set after assignment of random identification numbers to each entry and their encryption. Because noncoding intervening sequences (introns) are underrepresented in GenBank, a program that introduced (simulated) introns into mRNA and prokaryotic sequences was devised. In a further attempt to make the problem of identification more realistic, random base substitutions and single-base deletions were also incorporated. The randomly ordered entries were concatenated, along with random intergenic flanking sequences, into a single long "chromosome" 33 Mb in length and then cut into "cosmids" 50-100 kb long. The chopping process was conducted in such a way that terminal overlaps would allow the order of the entries in the chromosome to be reconstituted. Finally, the sequences of a substantial fraction of the cosmids were converted to their complements. Preliminary searching of 10 test cosmids revealed that more than two-thirds of the entries in the test set should be readily identifiable by type of gene product solely on the basis of comparison with the reference set. These preliminary results suggest that existing computer regimens and sequence collections would be able to identify the majority of eukaryotic genes in any new raw data set, the existence of introns not withstanding. Moreover, the analysis can be conducted in pace with the data collection so that the search results and summary identifications will be instantly available to the research community at large.

Animals↗

Creation of a large-scale genetic data bank for cardiovascular association studies.

BACKGROUND: A prerequisite for undertaking genetic association studies is the need for a genetic data bank with adequate DNA samples and a well-described clinical cohort. METHODS: We initiated a prospective single-center study enrolling 6,273 patients referred for cardiac catheterization in a genetic data bank (with eventual goal of 10,000 enrollees). Using a prescreening tool, the patients had comprehensive clinical phenotyping, including angiogram, electrocardiogram, echocardiogram, clinical history, and medication profile (Appendix A). Along with this clinical information, DNA, serum, plasma, basic metabolic panel, inflammation, and lipid panel were collected and stored in the database. RESULTS: Mean age of the patients enrolled was 64 +/- 12 years; 69% are men, 26% have diabetes, 79% have dyslipidemia, and 72% have coronary artery disease (CAD) > or = 50%. We undertook extensive quality-control measures to ensure the validity of both the clinical and DNA samples acquired into our GenBank. As part of this validation, we undertook a genetic association study to discern the effect of the apoE4 polymorphism on the risk for atherosclerosis. We are able to show that the apoE4 polymorphism is an independent risk factor for CAD. CONCLUSIONS: We have been able to create a large-scale genetic data bank as a resource to undertake genetic association studies. Key elements in implementation of this GenBank and baseline characteristics of our patient cohort are summarized. Lastly, as a "proof of concept" for the utility of this resource to discern gene variants associated with disease, we validated apoE4 polymorphism as an independent risk factor for CAD.

Aged↗

Influenza A H5N1 hemagglutinin cleavable signal sequence substitutions.

Eleven influenza A H5N1 hemagglutinin N-terminal cleavable signal sequences, coded by single nucleotide substitutions relative to reference A/Viet Nam/1203/2004, were identified by BLASTN search of GenBank and were characterized by molecular modeling. The signal sequences statistically segregated into two classes of states. Members of one class were uncharged and conformationally compact while members of the second class each carried a 2+ electric charge and were conformationally extended. Virtual signal sequences, not found on GenBank and based upon hypothetical transversions in the third codon, had molecular characteristics intermediate to those of the two classes of actual signal sequences. The high incidence of non-synonymous substitutions (63.6%), the high transition/transversion ratio (10/1) and the results of molecular modeling all suggest that the N-terminal cleavable signal sequence is mutationally evolving more rapidly than proteins which must assume specific conformational states in the mature influenza virion.

Evolution, Molecular↗

Dealing with repetitions in sequencing by hybridization.

DNA sequencing by hybridization (SBH) induces errors in the biochemical experiment. Some of them are random and disappear when the experiment is repeated. Others are systematic, involving repetitions in the probes of the target sequence. A good method for solving SBH problems must deal with both types of errors. In this work we propose a new hybrid genetic algorithm for isothermic and standard sequencing that incorporates the concept of structured combinations. The algorithm is then compared with other methods designed for handling errors that arise in standard and isothermic SBH approaches. DNA sequences used for testing are taken from GenBank. The set of instances for testing was divided into two groups. The first group consisted of sequences containing positive and negative errors in the spectrum, at a rate of up to 20%, excluding errors coming from repetitions. The second group consisted of sequences containing repeated oligonucleotides, and containing additional errors up to 5% added into the spectra. Our new method outperforms the best alternative procedures for both data sets. Moreover, the method produces solutions exhibiting extremely high degree of similarity to the target sequences in the cases without repetitions, which is an important outcome for biologists. The spectra prepared from the sequences taken from GenBank are available on our website http://bio.cs.put.poznan.pl/.

Algorithms↗

The novel organization and complete sequence of the ribosomal RNA gene of Nosema bombycis.

We present here for the first time the complete DNA sequence data (4301bp) of the ribosomal RNA (rRNA) gene of the microsporidian type species, Nosema bombycis. Sequences for the large subunit gene (LSUrRNA: 2497bp, GenBank Accession No. ), the internal transcribed spacer (ITS: 179bp, GenBank Accession No. ), the small subunit gene (SSUrRNA: 1232bp), intergenic spacer (IGS: 279bp), and 5S region (114bp) are also given, and the secondary structure of the large subunit is discussed. The organization of the N. bombycis rRNA gene is LSUrRNA-ITS-SSUrRNA-IGS-5S. This novel arrangement, in which the LSU is 5' of the SSU, is the reverse of the organizational sequence (i.e., SSU-ITS-LSU) found in all previously reported microsporidian rRNAs, including Nosema apis. This unique character in the type species may have taxonomic implications for the members of the genus Nosema.

Animals↗

Species identification using sequences of the trnL intron and the trnL-trnF IGS of chloroplast genome among popular plants in Taiwan.

Forensic botanical comparison can be hampered by the lack of appropriate DNA databases. While DNA sequence databases for many mitochondrial loci have been established for the identification of animal species, less is known regarding the genomes of plants. We report on the use of the trnL intron and the trnL-trnF intergenic spacer (IGS) in the chloroplast genome and establish a DNA sequence database for plant species identification. The DNA sequences at these two loci from commonly encountered plants, including monocots and dicots, were aligned to establish a DNA database of local plants. The database comprises 373 individual sequences representing 80 families, 206 genera and 269 species. These plant species can be grouped to species level using both sequence and length polymorphisms at these loci. To validate the database for future forensic purposes, we sequenced 20 blind samples and searched the local database and the databases of GenBank and EMBL. Fifteen of these 20 samples used in blind trial testing matched their respective species from our local DNA database but only 6 matched species registered in the GenBank and EMBL databases. The sequences of two species used in the blind trial did not match any sequence registered in any of these databases. Cluster analysis was performed to demonstrate the family and genus distribution of samples. Neighbor-joining trees of the two DNA regions from 70 samples of the local database and 10 of the species used in the blind trials were constructed and clustered to both family and genus. The bootstrap values of the trnL intron were higher than most of those of the trnL-trnF IGS. The sequence database described in this study can be used to identify plant species using DNA sequences of the trnL intron and trnL-trnF IGS of chloroplast genome and illustrates its value in plant species identification.

Cluster Analysis↗