Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequencing Resource”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Rapid sequence analysis of gene trap integrations to generate a resource of insertional mutations in mice.

Gene trapping in murine embryonic stem cells is a proven method for the simultaneous identification and mutation of genes in the mouse. Gene trap vectors are designed to detect insertions within genes through the production of a fusion mRNA transcript, making the identification of the endogenous gene possible by 5' rapid amplification of cDNA ends (RACE). Although the amplification of specific cDNAs can be achieved rapidly, cloning and screening of informative-sized cDNAs has proven to be time consuming. To eliminate the need for cloning, we have developed a method for solid-phase sequencing of 5' RACE products. More than 150 independent gene trap cell lines were analyzed, and sequence information was obtained for every line successfully amplified by RACE. With the vector used in this study, 40% of the cell lines were found to contain properly spliced gene trap events. The remaining lines were either spliced inefficiently or contained deletions of the vector. These results highlight the advantage of sequencing gene trap integrations before further characterization. This work now paves the way for large-scale gene trap screens in mice and should greatly accelerate the functional analysis of the mammalian genome.

Animals↗

Construction of a BAC library of Korean ginseng and initial analysis of BAC-end sequences.

We estimated the genome size of Korean ginseng (Panax ginseng C.A. Meyer), a medicinal herb, constructed a HindIII BAC library, and analyzed BAC-end sequences to provide an initial characterization of the library. The 1C nuclear DNA content of Korean ginseng was estimated to be 3.33 pg (3.12 x 10(3) Mb). The BAC library consists of 106,368 clones with an average size of 98.61 kb, amounting to 3.34 genome equivalents. Sequencing of 2167 BAC clones generated 2492 BAC-end sequences with an average length of 400 bp. Analysis using BLAST and motif searches revealed that 10.2%, 20.9% and 3.8% of the BAC-end sequences contained protein-coding regions, transposable elements and microsatellites, respectively. A comparison of the functional categories represented by the protein-coding regions found in BAC-end sequences with those of Arabidopsis revealed that proteins pertaining to energy metabolism, subcellular localization, cofactor requirement and transport facilitation were more highly represented in the P. ginseng sample. In addition, a sequence encoding a glucosyltransferase-like protein implicated in the ginsenoside biosynthesis pathway was also found. The majority of the transposable element sequences found belonged to the gypsy type (67.6%), followed by copia (11.7%) and LINE (8.0%) retrotransposons, whereas DNA transposons accounted for only 2.1% of the total in our sequence sample. Higher levels of transposable elements than protein-coding regions suggest that mobile elements have played an important role in the evolution of the genome of Korean ginseng, and contributed significantly to its complexity. We also identified 103 microsatellites with 3-38 repeats in their motifs. The BAC library and BAC-end sequences will serve as a useful resource for physical mapping, positional cloning and genome sequencing of P. ginseng.

Chromosomes, Artificial, Bacterial↗

Analysis of mercuric reductase (merA) gene diversity in an anaerobic mercury-contaminated sediment enrichment.

The reduction of ionic mercury to elemental mercury by the mercuric reductase (MerA) enzyme plays an important role in the biogeochemical cycling of mercury in contaminated environments by partitioning mercury to the atmosphere. This activity, common in aerobic environments, has rarely been examined in anoxic sediments where production of highly toxic methylmercury occurs. Novel degenerate PCR primers were developed which span the known diversity of merA genes in Gram-negative bacteria and amplify a 285 bp fragment at the 3' end of merA. These primers were used to create a clone library and to analyse merA diversity in an anaerobic sediment enrichment collected from a mercury-contaminated site in the Meadowlands, New Jersey. A total of 174 sequences were analysed, representing 71 merA phylotypes and four novel MerA clades. This first examination of merA diversity in anoxic environments suggests an untapped resource for novel merA sequences.

Anaerobiosis↗

Deletion and insertion in vivo somatic mutations in the hypoxanthine phosphoribosyltransferase (hprt) gene of human T-lymphocytes.

Deletion and insertion mutations have been found to be a major component of the in vivo somatic mutation spectrum in the hypoxanthine phosphoribosyltransferase (hprt) gene of T-lymphocytes. In a population of 172 healthy people (average age, 34; mutant frequency, 10.3 x 10(-6)), deletion/insertion mutations constituted 41% (89) of the 217 independent mutations, the remainder being base substitutions. Mutations were identified by multiplex PCR assay of genomic DNA for exon regions, by sequencing cDNA, or sequencing genomic DNA. The deletion and insertion mutations were divided among +/- 1 to 2 basepair (bp) frameshifts (14%, 30), small deletions and insertions of 3-200 bps (13%, 28), large deletions of one or more exons (12%, 27), and complex events (2%, 4). Frameshift mutations were dominated by -1 bp deletions (21 of 30). Exon 3 contained five frameshift mutations in the run of 6 Gs, the only site in the coding region with multiple frameshift mutations, possibly caused by strand dislocation during replication. Both endpoints were sequenced for 23 of the 28 small deletions/insertions including two tandem duplication events in exon 6. More small deletions (8/28), possibly mediated by trinucleotide repeats, occurred in exon 2 than in the other exons. Large deletions included total gene deletions (6), exon 2 + 3 deletions (4), and loss of multiple (9) and single exons (8) in genomic DNA. The diverse mutation spectrum indicates that multiple mechanisms operated at many different sequences and provides a resource for examination of deletion mutation.

Adult↗

AMPDB: the Arabidopsis Mitochondrial Protein Database.

The Arabidopsis Mitochondrial Protein Database is an Internet-accessible relational database containing information on the predicted and experimentally confirmed protein complement of mitochondria from the model plant Arabidopsis thaliana (http://www.ampdb.bcs.uwa.edu.au/). The database was formed using the total non-redundant nuclear and organelle encoded sets of protein sequences and allows relational searching of published proteomic analyses of Arabidopsis mitochondrial samples, a set of predictions from six independent subcellular-targeting prediction programs, and orthology predictions based on pairwise comparison of the Arabidopsis protein set with known yeast and human mitochondrial proteins and with the proteome of Rickettsia. A variety of precomputed physical-biochemical parameters are also searchable as well as a more detailed breakdown of mass spectral data produced from our proteomic analysis of Arabidopsis mitochondria. It contains hyperlinks to other Arabidopsis genomic resources (MIPS, TIGR and TAIR), which provide rapid access to changing gene models as well as hyperlinks to T-DNA insertion resources, Massively Parallel Signature Sequencing (MPSS) and Genome Tiling Array data and a variety of other Arabidopsis online resources. It also incorporates basic analysis tools built into the query structure such as a BLAST facility and tools for protein sequence alignments for convenient analysis of queried results.

Arabidopsis Proteins↗

Ultrastructural localization of giardins to the edges of disk microribbons of Giarida lamblia and the nucleotide and deduced protein sequence of alpha giardin.

The giardins are a group of 29-38-kD proteins in the ventral disk of the protozoan parasite Giardia lamblia. The disk attaches the parasite to the host's intestinal epithelium and is composed of parallel, coiled microtubules that are adjacent to the ventral plasma membrane and from which processes called microribbons extend into the cytoplasm; the microribbons are connected by crossbridges. G. lamblia cytoskeletons, consisting of disks and attached flagella, were isolated and used to show that the 29-38-kD proteins separate into five bands by one-dimensional electrophoresis and into 23 species by two-dimensional analysis. Rabbit antibodies raised against a 33-kD protein band, purified by one-dimensional gel electrophoresis and shown to contain three proteins by two-dimensional electrophoresis, recognized 17 proteins by two-dimensional immunoblot analysis. By immunofluorescence these antibodies reacted with the ventral disk but not with the flagella in isolated cytoskeletons. Electron microscopy revealed that the anti-giardin antibodies bound to the edges of the microribbons but not to the microtubules, crossbridges, or other, nondisk structures. Antibodies to tubulin reacted with both the disk and flagella in isolated cytoskeletons but bound only to the microtubules in these structures. The amino-terminal sequence of the 33-kD immunogen was determined and used to construct a DNA oligomer, and the oligomer was used to isolate the alpha giardin gene. The gene was used to hybrid select RNA, and the in vitro translation product from this RNA was precipitated by the antibodies against the 33-kD immunogen. The gene sequence was a single open reading frame of 885 nucleotides that predicted a protein of 33.8 kD. The protein sequence is unique, having no significant homology to two other giardin sequences or to any sequences within the Protein Identification Resource. It is predicted to be 82% alpha helical. The downstream sequence of the gene indicates that the sequence AGT-PuAA is located six to nine nucleotides beyond the stop codon in all protein-encoding genes of G. lamblia that have been sequenced and reported to date.

Amino Acid Sequence↗

The tmRNA website.

The tmRNA Website collects all available tmRNA sequences into a single public resource, along with alignments and a guide to searching for new sequences. Over the last year, several sequences have been updated or newly found by monitoring ongoing genome sequencing projects; tmRNA sequence data from 70 species are now available. New features include: color-coding of sequences to mark suggested base-paired regions, a list of the literature concerning tmRNA, careful crediting of tmRNA sequence identifications, and a split browser window. Updates are very frequent. The tmRNA Website has a new URL: http:www.indiana.edu/tmrna

Internet↗

Identification of a brain-specific human cerebrospinal fluid glycoprotein, beta-trace protein.

A prominent human cerebrospinal fluid (CSF) protein, P5, identified at mass 19-24 kDa and charge 5.5, by two-dimensional electrophoresis (2DE) and silver staining, has been previously demonstrated to be reduced in quantity in the CSF of patients with multiple sclerosis and schizophrenia. We report the purification and partial amino acid sequences from five tryptic fragments of P5. These sequences are not those of any known sequence in the Protein Identification Resource (PIR release 31) database. Synthetic peptides from two of the sequences were used to raise rabbit polyclonal antibodies. These antibodies detected P5 on 2DE blots of normal CSF proteins and other proteins of the same mass with a charge distribution between 5.17-8.5. These proteins comprise 5-10% of the total CSF protein and their mass, charge, abundance and predominance in CSF over plasma are consistent with a protein that had been initially characterized with antibodies, beta-trace protein. Glycosidase studies confirm that most of these proteins are due to sialic acid modifications that are N-linked to an 18 kDa protein, but other charge and mass variations also exist. 2DE blots of 26 types of human tissue and body fluid were immunostained. Of these, anti-P5 serum detected proteins of the same mass and charge as beta-trace protein only in brain samples. Proteins of different mass and charge from beta-trace protein were clearly immunostained in samples of eight tissues.(ABSTRACT TRUNCATED AT 250 WORDS)

Adult↗

SNP@Domain: a web resource of single nucleotide polymorphisms (SNPs) within protein domain structures and sequences.

The single nucleotide polymorphisms (SNPs) in conserved protein regions have been thought to be strong candidates that alter protein functions. Thus, we have developed SNP@Domain, a web resource, to identify SNPs within human protein domains. We annotated SNPs from dbSNP with protein structure-based as well as sequence-based domains: (i) structure-based using SCOP and (ii) sequence-based using Pfam to avoid conflicts from two domain assignment methodologies. Users can investigate SNPs within protein domains with 2D and 3D maps. We expect this visual annotation of SNPs within protein domains will help scientists select and interpret SNPs associated with diseases. A web interface for the SNP@Domain is freely available at http://snpnavigator.net/ and from http://bioportal.net/.

Computer Graphics↗

Computational analysis of protein tyrosine phosphatases: practical guide to bioinformatics and data resources.

The exponential growth of sequence data has become a challenge to database curators and end-users alike and biologists seeking to utilize the data effectively are faced with numerous analysis methods. Here, with practical examples from our bioinformatics analysis of the protein tyrosine phosphatases (PTPs), we show how computational analysis can be exploited to fuel hypothesis-driven experimental research through the exploration of online databases. We cover the following elements: (i) similarity searches and strategies to collect a non-redundant database of tyrosine-specific PTP domains; (ii) utilization of this database to classify human, fly, and worm PTPs (based on alignments and phylogenetic analysis); (iii) three-dimensional structural analysis to identify conserved regions (structure-function) and non-conserved selectivity-determining regions (substrate specificity); and (iv) genomic analysis, including mapping of exon structure, identification of pseudogenes, and exploration of disease databases. We discuss the importance of manual curation, illustrating examples in which pseudogenes give rise to predicted proteins in GenBank and note that domain servers, such as PFAM and SMART, erroneously include dual-specificity and lipid phosphatases in their collection of tyrosine-specific PTPs. To capitalize on our annotated set of 402 PTP domains (from 47 species and five phyla), we identify sequence conservation across taxonomic categories and explore structure-function relationships among tandem domain receptor-like PTPs. We define three Src homology 2 domain-containing PTP genes in stingray, zebrafish, and fugu and speculate on their evolutionary relationship with human pseudogenes. Our annotated sequences, along with a web service for phylogenetic classification of PTP domains, are available online (http://ptp.cshl.edu and http://science.novonordisk.com/ptp).

Amino Acid Sequence↗

PineappleDB: an online pineapple bioinformatics resource.

BACKGROUND: A world first pineapple EST sequencing program has been undertaken to investigate genes expressed during non-climacteric fruit ripening and the nematode-plant interaction during root infection. Very little is known of how non-climacteric fruit ripening is controlled or of the molecular basis of the nematode-plant interaction. PineappleDB was developed to provide the research community with access to a curated bioinformatics resource housing the fruit, root and nematode infected gall expressed sequences. DESCRIPTION: PineappleDB is an online, curated database providing integrated access to annotated expressed sequence tag (EST) data for cDNA clones isolated from pineapple fruit, root, and nematode infected root gall vascular cylinder tissues. The database currently houses over 5600 EST sequences, 3383 contig consensus sequences, and associated bioinformatic data including splice variants, Arabidopsis homologues, both MIPS based and Gene Ontology functional classifications, and clone distributions. The online resource can be searched by text or by BLAST sequence homology. The data outputs provide comprehensive sequence, bioinformatic and functional classification information. CONCLUSION: The online pineapple bioinformatic resource provides the research community with access to pineapple fruit and root/gall sequence and bioinformatic data in a user-friendly format. The search tools enable efficient data mining and present a wide spectrum of bioinformatic and functional classification information. PineappleDB will be of broad appeal to researchers investigating pineapple genetics, non-climacteric fruit ripening, root-knot nematode infection, crassulacean acid metabolism and alternative RNA splicing in plants.

Alternative Splicing↗

Worming your way through the genome.

The 100 Mb sequence of the nematode Caenorhabditis elegans genome will be completed in 1998. More than 10,000 predicted genes have been identified to date, so it should come as no surprise to find a C. elegans homologue of your favourite gene in current databases. For some investigators, the discovery of a C. elegans homologue represents a unique opportunity to adopt a genetic approach and to take advantage of the extensive repertoire of C. elegans gene characterization and manipulation tools. RNA injection provides a quick and efficient method for obtaining clues about wild-type gene function. Reverse genetic approaches also make it feasible to screen de novo for mutations in specific gene sequences. This review highlights the resources available for analysing a C. elegans homologue, starting from the gene sequence and proceeding to the biological function.

Animals↗

NCBI genetic resources supporting immunogenetic research.

The NCBI creates and maintains a set of integrated bibliographic, sequence, map, structure and other database resources to promote the efficient retrieval of information and the discovery of novel relationships. The connections made between elements of these resources permit researchers to start a search from a wide spectrum of entry points. These multiple dimensions of data can be roughly categorized by primary content as text or bibliographic (PubMed, PubMedCentral, OMIM, LocusLink), sequence (GenBank, Reference Sequence Project (RefSeq), dbSNP, MMDB), protein structure (MMDB) or map position (MapView). They can also becategorized by level of expert curation, which may range from validation of submissions from external groups (GenBank, PubMed, PubMedCentral,), to automatic computation (HomoloGene, UniGene), and to highly reviewed and corrected (LocusLink, MMDB, OMIM, RefSeq). Searches can be made by words (in an article title, key words, sequence annotation, database value, author) by sequence (BLAST or e-PCR against multiple sequence databases), or by map coordinates. By computing or curating bi-directional links between related objects, NCBI can represent content on the genetics, molecular biology, and clinical considerations of interest to immunogeneticists. There is also an emerging resource developed by the NCBI in collaboration with the IHWG devoted to the presentation of MHC data (dbMHC). How dbMHC will augment existing resources at the NCBI is described.

Computational Biology↗

Frequency, type, distribution and annotation of simple sequence repeats in Rosaceae ESTs.

Genomic resources for peach, a model species for Rosaceae, are being developed to accelerate gene discovery in other Rosaceae species by comparative mapping. Simple sequence repeats (SSRs) are an important tool for comparative mapping because of their high polymorphism and transportability. To accelerate the development of SSR markers, we analyzed publicly available Rosaceae expressed sequence tags (ESTs) for SSRs. A total of 17,284 ESTs from almond, peach and rose were assembled into putatively non-redundant EST sets. For comparison, 179,099 ESTs from Arabidopsis were also used in the analysis. About 4% of the assembled ESTs contained SSRs in Rosaceae, which was higher than the 2.4% found in Arabidopsis. About half of the SSRs were found in the putative UTR, and the estimated average distance between SSRs in the UTR was 5.5 kb in rose, 5.1 kb in almond, 7 kb in peach and 13 kb in Arabidopsis. In the putative coding region, the estimated average distance was two to four times longer than in the UTR. Rosaceae ESTs containing SSRs were functionally annotated using the GenBank nr database and further classified using the gene ontology terms associated with the matching sequences in the SwissProt database. The detailed data including the sequences and annotation results are available from http://www.genome.clemson.edu/gdr/rosaceaessr/.

Arabidopsis↗

Database of protein sequence alignments: PIR-ALN.

The Protein Information Resource (PIR) has been maintaining a database of curated protein sequence alignments since 1991. The collection includes superfamily, family and homology domain alignments. CLUSTAL V/W is used to generate multiple sequence alignments and ALNED, an interactive alignment editor, is used to check and correct them. The database has helped in classifying sequences, in defining new homology domains, and in spreading and standardizing protein names, features and keywords among members of a family or superfamily. The ATLAS information retrieval system can be used to browse and query the PIR-ALN alignments. The quarterly and weekly updates can be accessed via the WWW at http://www-nbrf. georgetown.edu/pir/

Databases, Factual↗

Bases and spaces: resources on the web for accessing the draft human genome.

SUMMARY: Much is expected of the draft human genome sequence, and yet there is no central resource to host the plethora of sequence and mapping information available. Consequently, finding the most useful and reliable human genome data and resources currently available on the web can be challenging, but is not impossible.

Cloning, Molecular↗

Analysis of T-DNA insertion site distribution patterns in Arabidopsis thaliana reveals special features of genes without insertions.

Large collections of sequence-indexed T-DNA insertion mutants are invaluable resources for plant functional genomics. Flanking sequence tag (FST) data from these collections indicated that T-DNA insertions are not randomly distributed in the Arabidopsis thaliana genome and that there are still a fairly high number of annotated genes without T-DNA insertions. We have analyzed FST data from the FLAGdb, GABI-Kat, and SIGnAL mutant populations. The lack of detectable transcriptional activity and the absence of suitable restriction sites were among the reasons genes are not covered by insertions. Additionally, a refined analysis of FSTs to genes with annotated noncoding regions showed that transcription initiation and polyadenylation site regions of genes are favored targets for T-DNA integration. These findings have implications for the use of T-DNA in saturation mutagenesis and for our chances to find a useful knockout allele for every gene.

Arabidopsis↗

A time to sequence.

Explore the source record for details and available documents.

Chromosome Mapping↗