Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequencing Resource”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

The Protein Information Resource (PIR) and the PIR-International Protein Sequence Database.

From its origin, the PIR has aspired to support research in computational biology and genomics through the compilation of a comprehensive, quality controlled and well-organized protein sequence information resource. The resource originated with the pioneering work of the late Margaret O. Dayhoff in the early 1960s. Since 1988, the Protein Sequence Database has been maintained collaboratively by PIR-International, an association of macromolecular sequence data collection centers dedicated to fostering international cooperation as an essential element in the development of scientific databases. The work of the resource is widely distributed and is available on the World Wide Web, via FTP, E-mail server, CD-ROM and magnetic media. It is widely redistributed and incorporated into many other protein sequence data compilations including SWISS-PROT and theEntrezsystem of the NCBI.

Amino Acid Sequence↗

A web-based genotyping resource for viral sequences.

The Genotyping tool at the National Center for Biotechnology Information is a web-based program that identifies the genotype (or subtype) of recombinant or non-recombinant viral nucleotide sequences. It works by using BLAST to compare a query sequence to a set of reference sequences for known genotypes. Predefined reference genotypes exist for three major viral pathogens: human immunodeficiency virus 1 (HIV-1), hepatitis C virus (HCV) and hepatitis B virus (HBV). User-defined reference sequences can be used at the same time. The query sequence is broken into segments for comparison to the reference so that the mosaic organization of recombinant sequences could be revealed. The results are displayed graphically using color-coded genotypes. Therefore, the genotype(s) of any portion of the query can quickly be determined. The Genotyping tool can be found at: http://www.ncbi.nih.gov/projects/genotyping/formpage.cgi.

Algorithms↗

SeqHelp: a program to analyze molecular sequences utilizing common computational resources.

Here we describe a tool to analyze molecular sequences utilizing the internet and existing computational resources for molecular biology. The computer program SeqHelp organizes information from database searches, gene structure prediction, and other information to generate multiply aligned, hypertext-linked reports to allow for fast analysis of molecular sequences. The efficient and economical strategy in this program can be employed to study molecular sequences for gene cloning, mutation analysis, and identical sequence search projects.

Algorithms↗

Seq2Struct: a resource for establishing sequence-structure links.

UNLABELLED: Several methods for establishing cross-links between Protein Data Bank (PDB) structures or Structural Classification of Proteins (SCOP) domains and Swiss-Prot + TrEMBL sequences (or vice versa) rely on database annotations. Alternatively, sequence alignment procedures can be used. In this study, we describe Seq2Struct, a web resource for the identification of sequence-structure links. The resource consists of an exhaustive collection of annotated links between Swiss-Prot + TrEMBL and PDB + SCOP database entries. Links are based on pre-established highly reliable thresholds and stored in a relational database, which has been enhanced using annotations derived from Swiss-Prot, PDB, SCOP, GOA and DSSP databases. The Seq2Struct database contents, supported by a WWW web interface, can be queried both online and downloaded. AVAILABILITY: The Seq2Struct resource, with related documentation, is available at http://surface.bio.uniroma2.it/seq2struct/ CONTACT: seq2struct@cbm.bio.uniroma2.it.

Database Management Systems↗

Regulatory sequence analysis tools.

The web resource Regulatory Sequence Analysis Tools (RSAT) (http://rsat.ulb.ac.be/rsat) offers a collection of software tools dedicated to the prediction of regulatory sites in non-coding DNA sequences. These tools include sequence retrieval, pattern discovery, pattern matching, genome-scale pattern matching, feature-map drawing, random sequence generation and other utilities. Alternative formats are supported for the representation of regulatory motifs (strings or position-specific scoring matrices) and several algorithms are proposed for pattern discovery. RSAT currently holds >100 fully sequenced genomes and these data are regularly updated from GenBank.

5' Flanking Region↗

Analysis of multiple genomic sequence alignments: a web resource, online tools, and lessons learned from analysis of mammalian SCL loci.

Comparative analysis of genomic sequences is becoming a standard technique for studying gene regulation. However, only a limited number of tools are currently available for the analysis of multiple genomic sequences. An extensive data set for the testing and training of such tools is provided by the SCL gene locus. Here we have expanded the data set to eight vertebrate species by sequencing the dog SCL locus and by annotating the dog and rat SCL loci. To provide a resource for the bioinformatics community, all SCL sequences and functional annotations, comprising a collation of the extensive experimental evidence pertaining to SCL regulation, have been made available via a Web server. A Web interface to new tools specifically designed for the display and analysis of multiple sequence alignments was also implemented. The unique SCL data set and new sequence comparison tools allowed us to perform a rigorous examination of the true benefits of multiple sequence comparisons. We demonstrate that multiple sequence alignments are, overall, superior to pairwise alignments for identification of mammalian regulatory regions. In the search for individual transcription factor binding sites, multiple alignments markedly increase the signal-to-noise ratio compared to pairwise alignments.

Animals↗

Arabidopsis genomic information for interpreting wheat EST sequences.

The resources available from Arabidopsis thaliana for interpreting functional attributes of wheat EST are reviewed. A focus for the review is a comparison between wheat EST sequences, generated from developing endosperm tissue, and the complete genomic sequence from Arabidopsis. The available information indicates that not only can tentative annotations be assigned to many wheat genes but also putative or unknown Arabidopsis gene annotations can be improved by comparative genomics.

Amino Acid Sequence↗

Screening for novel enzymes for biocatalytic processes: accessing the metagenome as a resource of novel functional sequence space.

Historically, biotechnology has missed up to 99% of existing microbial resources by using traditional screening techniques. Strategies of directly cloning 'environmental DNA' comprising the genetic blueprints of entire microbial consortia (the so-called 'metagenome') provide molecular sequence space that along with ingenious in vitro evolution technologies will act synergistically to bring a maximum of available sequence-space into biocatalytic application.

Bacteria↗

PatGen--a consolidated resource for searching genetic patent sequences.

UNLABELLED: Compared to the wealth of online resources covering genomic, proteomic and derived data the Bioinformatics community is rather underserved when it comes to patent information related to biological sequences. The current online resources are either incomplete or rather expensive. This paper describes, PatGen, an integrated database containing data from bioinformatic and patent resources. This effort addresses the inconsistency of publicly available genetic patent data coverage by providing access to a consolidated dataset. AVAILABILITY: PatGen can be searched at http://www.patgendb.com CONTACT: rjdrouse@patentinformatics.com.

Abstracting and Indexing↗

The TIGR Plant Repeat Databases: a collective resource for the identification of repetitive sequences in plants.

In a number of higher plants, a substantial portion of the genome is composed of repetitive sequences that can hinder genome annotation and sequencing efforts. To better understand the nature of repetitive sequences in plants and provide a resource for identifying such sequences, we constructed databases of repetitive sequences for 12 plant genera: Arabidopsis, Brassica, Glycine, Hordeum, Lotus, Lycopersicon, Medicago, Oryza, Solanum, Sorghum, Triticum and Zea (www.tigr.org/tdb/e2k1/plant. repeats/index.shtml). The repetitive sequences within each database have been coded into super-classes, classes and sub-classes based on sequence and structure similarity. These databases are available for sequence similarity searches as well as downloadable files either as entire databases or subsets of each database. To further the utility for comparative studies and to provide a resource for searching for repetitive sequences in other genera within these families, repetitive sequences have been combined into four databases to represent the Brassicaceae, Fabaceae, Gramineae and Solanaceae families. Collectively, these databases provide a resource for the identification, classification and analysis of repetitive sequences in plants.

Computational Biology↗

PathoGene: a pathogen coding sequence discovery and analysis resource.

PathoGene is a web-based resource that streamlines the process of predicting genes in microorganisms and designs PCR primers for amplification to facilitate sequence analysis and experimentation. PathoGene currently supports primer design for every complete microbial, viral, and fungal genome as cataloged in GenBank by the National Center for Biotechnology Information (NCBI; http://www.ncbi.nlm.nih.gov/). The resulting primers can then be subjected to a stand-alone Basic Local Alignment Search Tool (BLAST) system called PathoBLAST in which the predicted PCR product and/or primers can be compared against the genome of interest or a similar genome to find related genes or estimate primer quality.

Bacillus anthracis↗

PFDB: a generic protein family database integrating the CATH domain structure database with sequence based protein family resources.

MOTIVATION: The PFDB (Protein Family Database) is a new database designed to integrate protein family-related data with relevant functional and genomic data. It currently manages biological data for three projects-the CATH protein domain database (Orengo et al., 1997; Pearl et al., 2001), the VIDA virus domains database (Albà et al., 2001) and the Gene3D database (Buchan et al., 2001). The PFDB has been designed to accommodate protein families identified by a variety of sequence based or structure based protocols and provides a generic resource for biological research by enabling mapping between different protein families and diverse biochemical and genetic data, including complete genomes. RESULTS: A characteristic feature of the PFDB is that it has a number of meta-level entities (for example aggregation, collection and inclusion) represented as base tables in the final design. The explicit representation of relationships at the meta-level has a number of advantages, including flexibility-both in terms of the range of queries that can be formulated and the ability to integrate new biological entities within the existing design. A potential drawback with this approach-poor performance caused by the number of joins across meta-level tables-is avoided by implementing the PFDB with materialized views using the mature relational database technology of Oracle 8i. The resultant database is both fast and flexible. This paper presents the principles on which the database has been designed and implemented, and describes the current status of the database and query facilities supported.

Database Management Systems↗

The SDH mutation database: an online resource for succinate dehydrogenase sequence variants involved in pheochromocytoma, paraganglioma and mitochondrial complex II deficiency.

BACKGROUND: The SDHA, SDHB, SDHC and SDHD genes encode the subunits of succinate dehydrogenase (succinate: ubiquinone oxidoreductase), a component of both the Krebs cycle and the mitochondrial respiratory chain. SDHA, a flavoprotein and SDHB, an iron-sulfur protein together constitute the catalytic domain, while SDHC and SDHD encode membrane anchors that allow the complex to participate in the respiratory chain as complex II. Germline mutations of SDHD and SDHB are a major cause of the hereditary forms of the tumors paraganglioma and pheochromocytoma. The largest subunit, SDHA, is mutated in patients with Leigh syndrome and late-onset optic atrophy, but has not as yet been identified as a factor in hereditary cancer. DESCRIPTION: The SDH mutation database is based on the recently described Leiden Open (source) Variation Database (LOVD) system. The variants currently described in the database were extracted from the published literature and in some cases annotated to conform to current mutation nomenclature. Researchers can also directly submit new sequence variants online. Since the identification of SDHD, SDHC, and SDHB as classic tumor suppressor genes in 2000 and 2001, studies from research groups around the world have identified a total of 120 variants. Here we introduce all reported paraganglioma and pheochromocytoma related sequence variations in these genes, in addition to all reported mutations of SDHA. The database is now accessible online. CONCLUSION: The SDH mutation database offers a valuable tool and resource for clinicians involved in the treatment of patients with paraganglioma-pheochromocytoma, clinical geneticists needing an overview of current knowledge, and geneticists and other researchers needing a solid foundation for further exploration of both these tumor syndromes and SDHA-related phenotypes.

Codon, Nonsense↗

A genome-wide, end-sequenced 129Sv BAC library resource for targeting vector construction.

The majority of gene-targeting experiments in mice are performed in 129Sv-derived embryonic stem (ES) cell lines, which are generally considered to be more reliable at colonizing the germ line than ES cells derived from other strains. Gene targeting is reliant on homologous recombination of a targeting vector with the host ES cell genome. The efficiency of recombination is affected by many factors, including the isogenicity (H. te Riele et al., 1992, Proc. Natl. Acad. Sci. USA 89, 5128-5132) and the length of homologous sequence of the targeting vector and the location of the target locus. Here we describe the double-end sequencing and mapping of 84,507 bacterial artificial chromosomes (BACs) generated from AB2.2 ES cell DNA (129S7/SvEvBrd-Hprtb-m2). We have aligned these BACs against the mouse genome and displayed them on the Ensembl genome browser, DAS: 129S7/AB2.2. This library has an average insert size of 110.68 kb and average depth of genome coverage of 3.63- and 1.24-fold across the autosomes and sex chromosomes, respectively. Over 97% of the mouse genome and 99.1% of Ensembl genes are covered by clones from this library. This publicly available BAC resource can be used for the rapid construction of targeting vectors via recombineering. Furthermore, we show that targeting vectors containing DNA recombineered from this BAC library can be used to target genes efficiently in several 129-derived ES cell lines.

Animals↗

Human aldehyde dehydrogenase. cDNA cloning and primary structure of the enzyme that catalyzes dehydrogenation of 4-aminobutyraldehyde.

Human liver aldehyde dehydrogenase (E3 isozyme), with wide substrate specificity and low Km for 4-aminobutyraldehyde, was only recently characterized [Kurys, G., Ambroziak, W. & Pietruszko, R. (1989) J. Biol. Chem. 264, 4715-4721] and in this study we report on its primary structure. Polyclonal antibodies, specific for the E3 isozyme and three oligonucleotide probes derived from amino acid sequence of the E3 protein, were used for isolation of the first cDNA clone encoding the human enzyme (1503 bp; coding for 440 amino acid residues). Additional clones were obtained by using the first isolated clone as a probe. The largest clone of 1635 bp coded for 462 amino acid residues; it was longer at the 3'end of the cDNA non-coding region. The identity of the clone was established by DNA sequencing and by comparison with peptide sequences derived from the E3 protein, which constituted approximately 29% of the total primary structure of the E3 isozyme. The start codon was never encountered despite a variety of different approaches (500 amino acid residues were expected on the basis of SDS-gel molecular-mass determination of the E3 isozyme subunit). Despite the great catalytic similarity between the E3 and E1 isozymes [Ambroziak, W. & Pietruszko, R. (1991) J. Biol. Chem. 266, 13011-13018], the primary structure of the E3 isozyme has only approximately 40.6% of positional identity with that of the E1 isozyme. Sequence comparison with GenBank and Protein Identification Resource database sequences indicated no primary structure of aldehyde dehydrogenase more closely resembling the E3 isozyme than that of Escherichia coli betaine aldehyde dehydrogenase (52.7% positional identity), a prokaryotic enzyme specific for betaine aldehyde.

Aldehyde Dehydrogenase↗