Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequencing Resource”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 829 records · Page 46Linked to original sources

The Arabidopsis unannotated secreted peptide database, a resource for plant peptidomics.

In the era of genomics, if a gene is not annotated, it is not investigated. Due to their small size, genes encoding peptides are often missed in genome annotations. Secreted peptides are important regulators of plant growth, development, and physiology. Identification of additional peptide signals by sequence homology searches has had limited success due to sequence heterogeneity. A bioinformatics approach was taken to find unannotated Arabidopsis (Arabidopsis thaliana) peptides. Arabidopsis chromosome sequences were searched for all open reading frames (ORFs) encoding peptides and small proteins between 25 and 250 amino acids in length. The translated ORFs were then sequentially queried for the presence of an amino-terminal cleavable signal peptide, the absence of transmembrane domains, and the absence of endoplasmic reticulum lumenal retention sequences. Next, the ORFs were filtered against the The Arabidopsis Information Resource 6.0 annotated Arabidopsis genes to remove those ORFs overlapping known genes. The remaining 33,809 ORFs were placed in a relational database to which additional annotation data were deposited. Genome-wide tiling array data were compared with the coordinates of the ORFs, supporting the possibility that many of the ORFs may be expressed. In addition, clustering and sequence similarity analyses revealed that many of the putative peptides are in gene families and/or appear to be present in the rice (Oryza sativa) genome. A subset of the ORFs was evaluated by reverse transcription-PCR and, for one-fifth of those, expression was detected. These results support the idea that the number and diversity of plant peptides is broader than currently assumed. The peptides identified and their annotation data may be viewed or downloaded through a searchable Web interface at peptidome.missouri.edu.

Arabidopsis↗

The FLEXGene repository: exploiting the fruits of the genome projects by creating a needed resource to face the challenges of the post-genomic era.

Thanks to the results of the multiple completed and ongoing genome sequencing projects and to the newly available recombination-based cloning techniques, it is now possible to build gene repositories with no precedent in their composition, formatting, and potential. This new type of gene repository is necessary to address the challenges imposed by the post-genomic era, i.e., experimentation on a genome-wide scale. We are building the FLEXGene (Full Length EXpression-ready) repository. This unique resource will contain clones representing the complete ORFeome of different organisms, including Homo sapiens as well as several pathogens and model organisms. It will consist of a comprehensive, characterized (sequence-verified), and arrayed gene repository. This resource will allow full exploitation of the genomic information by enabling genome-wide scale experimentation at the level of functional/phenotypic assays as well as at the level of protein expression, purification, and analysis. Here we describe the rationale and construction of this resource and focus on the data obtained from the Saccharomyces cerevisiae project.

Animals↗

A unique set of 11,008 onion expressed sequence tags reveals expressed sequence and genomic differences between the monocot orders Asparagales and Poales.

Enormous genomic resources have been developed for plants in the monocot order Poales; however, it is not clear how representative the Poales are for the monocots as a whole. The Asparagales are a monophyletic order sister to the lineage carrying the Poales and possess economically important plants such as asparagus, garlic, and onion. To assess the genomic differences between the Asparagales and Poales, we generated 11,008 unique ESTs from a normalized cDNA library of onion. Sequence analyses of these ESTs revealed microsatellite markers, single nucleotide polymorphisms, and homologs of transposable elements. Mean nucleotide similarity between rice and the Asparagales was 78% across coding regions. Expressed sequence and genomic comparisons revealed strong differences between the Asparagales and Poales for codon usage and mean GC content, GC distribution, and relative GC content at each codon position, indicating that genomic characteristics are not uniform across the monocots. The Asparagales were more similar to eudicots than to the Poales for these genomic characteristics.

Cytosine↗

The gene trap resource: a treasure trove for hemopoiesis research.

The laboratory mouse is an invaluable tool for functional gene discovery because of its genetic malleability and a biological similarity to human systems that facilitates identification of human models of disease. A number of mutagenic technologies are being used to elucidate gene function in the mouse. Gene trapping is an insertional mutagenesis strategy that is being undertaken by multiple research groups, both academic and private, in an effort to introduce mutations across the mouse genome. Large-scale, publicly funded gene trap programs have been initiated in several countries with the International Gene Trap Consortium coordinating certain efforts and resources. We outline the methodology of mammalian gene trapping and how it can be used to identify genes expressed in both primitive and definitive blood cells and to discover hemopoietic regulator genes. Mouse mutants with hematopoietic phenotypes derived using gene trapping are described. The efforts of the large-scale gene trapping consortia have now led to the availability of libraries of mutagenized ES cell clones. The identity of the trapped locus in each of these clones can be identified by sequence-based searching via the world wide web. This resource provides an extraordinary tool for all researchers wishing to use mouse genetics to understand gene function.

Animals↗

MPact: the MIPS protein interaction resource on yeast.

In recent years, the Munich Information Center for Protein Sequences (MIPS) yeast protein-protein interaction (PPI) dataset has been used in numerous analyses of protein networks and has been called a gold standard because of its quality and comprehensiveness [H. Yu, N. M. Luscombe, H. X. Lu, X. Zhu, Y. Xia, J. D. Han, N. Bertin, S. Chung, M. Vidal and M. Gerstein (2004) Genome Res., 14, 1107-1118]. MPact and the yeast protein localization catalog provide information related to the proximity of proteins in yeast. Beside the integration of high-throughput data, information about experimental evidence for PPIs in the literature was compiled by experts adding up to 4300 distinct PPIs connecting 1500 proteins in yeast. As the interaction data is a complementary part of CYGD, interactive mapping of data on other integrated data types such as the functional classification catalog [A. Ruepp, A. Zollner, D. Maier, K. Albermann, J. Hani, M. Mokrejs, I. Tetko, U. Güldener, G. Mannhaupt, M. Münsterkötter and H. W. Mewes (2004) Nucleic Acids Res., 32, 5539-5545] is possible. A survey of signaling proteins and comparison with pathway data from KEGG demonstrates that based on these manually annotated data only an extensive overview of the complexity of this functional network can be obtained in yeast. The implementation of a web-based PPI-analysis tool allows analysis and visualization of protein interaction networks and facilitates integration of our curated data with high-throughput datasets. The complete dataset as well as user-defined sub-networks can be retrieved easily in the standardized PSI-MI format. The resource can be accessed through http://mips.gsf.de/genre/proj/mpact.

Databases, Protein↗

CDD: a conserved domain database for interactive domain family analysis.

The conserved domain database (CDD) is part of NCBI's Entrez database system and serves as a primary resource for the annotation of conserved domain footprints on protein sequences in Entrez. Entrez's global query interface can be accessed at http://www.ncbi.nlm.nih.gov/Entrez and will search CDD and many other databases. Domain annotation for proteins in Entrez has been pre-computed and is readily available in the form of 'Conserved Domain' links. Novel protein sequences can be scanned against CDD using the CD-Search service; this service searches databases of CDD-derived profile models with protein sequence queries using BLAST heuristics, at http://www.ncbi.nlm.nih.gov/Structure/cdd/wrpsb.cgi. Protein query sequences submitted to NCBI's protein BLAST search service are scanned for conserved domain signatures by default. The CDD collection contains models imported from Pfam, SMART and COG, as well as domain models curated at NCBI. NCBI curated models are organized into hierarchies of domains related by common descent. Here we report on the status of the curation effort and present a novel helper application, CDTree, which enables users of the CDD resource to examine curated hierarchies. More importantly, CDD and CDTree used in concert, serve as a powerful tool in protein classification, as they allow users to analyze protein sequences in the context of domain family hierarchies.

Amino Acid Sequence↗

Human protein reference database as a discovery resource for proteomics.

The rapid pace at which genomic and proteomic data is being generated necessitates the development of tools and resources for managing data that allow integration of information from disparate sources. The Human Protein Reference Database (http://www.hprd.org) is a web-based resource based on open source technologies for protein information about several aspects of human proteins including protein-protein interactions, post-translational modifications, enzyme-substrate relationships and disease associations. This information was derived manually by a critical reading of the published literature by expert biologists and through bioinformatics analyses of the protein sequence. This database will assist in biomedical discoveries by serving as a resource of genomic and proteomic information and providing an integrated view of sequence, structure, function and protein networks in health and disease.

Computational Biology↗

The genexpress IMAGE knowledge base of the human muscle transcriptome: a resource of structural, functional, and positional candidate genes for muscle physiology and pathologies.

Sequence, gene mapping, and expression data corresponding to 910 genes transcribed in human skeletal muscle have been integrated to form the muscle module of the Genexpress IMAGE Knowledge Base. Based on cDNA array hybridization, a set of 14 transcripts preferentially or specifically expressed in muscle have been selected and characterized in more detail: Their pattern of expression was confirmed by Northern blot analysis; their structure was further characterized by full-insert cDNA sequencing and cDNA extension; the map location of the corresponding genes was refined by radiation hybrid mapping. Five of the 14 selected genes appear as interesting positional and functional candidate genes to study in relation with muscle physiology and/or specific orphan muscular pathologies. One example is discussed in more detail. The expression profiling data and the associated Genexpress Index2 entries for the 910 genes and the detailed characterization of the 14 selected transcripts are available from a dedicated Web server at. The database has been organized to provide the users with a working space where they can find curated, annotated, integrated data for their genes of interest. Different navigation routes to exploit the resource are discussed.

Base Sequence↗

Hydrophobic cluster analysis: an efficient new way to compare and analyse amino acid sequences.

A new method for comparing and aligning protein sequences is described. This method, hydrophobic cluster analysis (HCA), relies upon a two-dimensional (2D) representation of the sequences. Hydrophobic clusters are determined in this 2D pattern and then used for the sequence comparisons. The method does not require powerful computer resources and can deal with distantly related proteins, even if no 3D data are available. This is illustrated in the present report by a comparison of human haemoglobin with leghaemoglobin, a comparison of the two domains of liver rhodanese (thiosulphate sulphurtransferase) and a comparison of plastocyanin and azurin.

Amino Acid Sequence↗

Cloning the shared components of complex DNA resources.

The complex and repetitive nature of mammalian genomes limits the ability of conventional molecular techniques to recover sequences of interest. Here we describe a rapid and simple procedure for the direct cloning of sequences which are coincident between DNA mixtures of whole genome complexity. The system, called end ligation coincident sequence cloning (EL-CSC), can enrich coincident DNA by greater than 10(6)-fold and overcomes problems associated with repetitive elements. Applying EL-CSC to various paired DNA resources enables the facile cloning of both genomic markers and novel genes. To demonstrate the power of the method we have i) selectively purified single copy sequences from a complete genome, and ii) isolated gene fragments from 260 kb of cloned genomic DNA.

Base Sequence↗

Phosphoproteomic Analysis of Cortical Tissue from Mice Lacking Both CaMKIIα and CaMKIIβ Identifies Novel In Vivo Substrates.

Ca2+/calmodulin-dependent protein kinase II (CaMKII) plays a critical role in calcium signaling. Several studies have shown that mice with single Camk2a or Camk2b gene knockouts are viable, yet exhibit distinct phenotypes, whereas the double knockout of both genes is lethal. These findings indicate that each gene can have distinct roles and that they also partially compensate for each other in yet unknown essential brain functions. In order to provide insight into potential novel CaMKII functions, we performed parallel phosphoproteomic analyses on nonstimulated cortex tissues from inducible Camk2a and Camk2b double knockout (Camk2af/f;Camk2bf/f;CAG-CreESR) mice and from wild type mice. A total of 5622 phosphorylated peptides derived from 2080 proteins were identified. Phosphorylation at serine/threonine residues in 130 proteins was downregulated in the double knockout mice, including residues in 113 proteins that have not previously been identified as potential CaMKII substrates. Comparison of amino acid sequences surrounding the downregulated phosphorylation residues provided new insights into the CaMKII-substrate consensus sequences in vivo. This data set provides an important resource for future studies examining novel roles for CaMKII in the brain.

Animals↗

A collection of sequenced and mapped Ds transposon insertion sites in Arabidopsis thaliana.

Insertional mutagenesis is a powerful tool for generating knockout mutations that facilitate associating biological functions with as yet uncharacterized open reading frames (ORFs) identified by genomic sequencing or represented in EST databases. We have generated a collection of Dissociation (Ds) transposon lines with insertions on all 5 Arabidopsis chromosomes. Here we report the insertion sites in 260 independent single-transposon lines, derived from four different Ds donor sites. We amplified and determined the genomic sequence flanking each transposon, then mapped its insertion site by identity of the flanking sequences to the corresponding sequence in the Arabidopsis genome database. This constitutes the largest collection of sequence-mapped Ds insertion sites unbiased by selection against the donor site. Insertion site clusters have been identified around three of the four donor sites on chromosomes 1 and 5, as well as near the nucleolus organizers on chromosomes 2 and 4. The distribution of insertions between ORFs and intergenic sequences is roughly proportional to the ratio of genic to intergenic sequence. Within ORFs, insertions cluster near the translational start codon, although we have not detected insertion site selectivity at the nucleotide sequence level. A searchable database of insertion site sequences for the 260 transposon insertion sites is available at http://sgio2.biotec.psu.edu/sr. This and other collections of Arabidopsis lines with sequence-identified transposon insertion sites are a valuable genetic resource for functional genomics studies because the transposon location is precisely known, the transposon can be remobilized to generate revertants, and the Ds insertion can be used to initiate further local mutagenesis.

Arabidopsis↗

Improved techniques for the identification of pseudogenes.

MOTIVATION: Pseudogenes are the remnants of genomic sequences of genes which are no longer functional. They are frequent in most eukaryotic genomes, and an important resource for comparative genomics. However, pseudogenes are often mis-annotated as functional genes in sequence databases. Current methods for identifying pseudogenes include methods which rely on the presence of stop codons and frameshifts, as well as methods based on the ratio of non-silent to silent nucleotide substitution rates (dN/dS). A recent survey concluded that 50% of human pseudogenes have no detectable truncation in their pseudo-coding regions, indicating that the former methods lack sensitivity. The latter methods have been used to find sets of genes enriched for pseudogenes, but are not specific enough to accurately separate pseudogenes from expressed genes. RESULTS: We introduce a program called pseudogene inference from loss of constraint (PSILC) which incorporates novel methods for separating pseudogenes from functional genes. The methods calculate the log-odds score that evolution along the final branch of the gene tree to the query gene has been according to the following constraints: A neutral nucleotide model compared to a Pfam domain encoding model (PSILC(nuc/dom)); A protein coding model compared to a Pfam domain encoding model (PSILC(prot/dom)). Using the manual annotation of human chromosome 6, we show that both these methods result in a more accurate classification of pseudogenes than dN/dS when a Pfam domain alignment is available. AVAILABILITY: PSILC is available from http://www.sanger.ac.uk/Software/PSILC

Algorithms↗

Single-nucleotide polymorphism (SNP) analysis in the ABC half-transporter ABCG2 (MXR/BCRP/ABCP1).

Variations in the amino acid sequence of ABC transporters have been shown to impact substrate specificity. We identified two acquired mutations in ABCG2, the ABC half-transporter overexpressed in mitoxantrone-resistant cell lines. These mutations confer differences in substrate specificity and suggest that naturally occurring variants could also affect substrate specificity. To search for the existence of single nucleotide polymorphisms (SNPs) in ABCG2, we sequenced 90 ethnically diverse DNAs from the Single Nucleotide Polymorphism Discovery Resource representing the spectrum of human genotypes. We identified 3 noncoding SNPs in the untranslated regions, 3 nonsynonymous and 2 synonymous SNPs in the coding region and 7 SNPs in the intron sequences adjacent to the sixteen ABCG2 exons. Nonsynonymous SNPs at nucleotide 238 (V12M; exon 2) and nucleotide 625 (Q141K; exon 5) showed a greater frequency of heterozygosity (22.2% and 10%) than the SNP at 2062 (D620N; exon 16). Heterozygous changes at nucleotide 238 are in linkage disequilibrium with an SNP observed 36 bases downstream from the end of exon 2. No polymorphism at amino acid 482 was identified to correspond to the R to G or R to T mutations previously found in two drug resistant cell lines. Among 23 drug resistant sublines for which sequence at position 482 was determined, no additional mutations were found. Heterozygosity at amino acid 12 allowed us to identify overexpression of a single allele in a subset of drug resistant cell lines, a feature that could be exploited clinically in evaluating the significance of ABCG2 expression in malignancy. We conclude that ABCG2 is well conserved and that described amino acid polymorphisms seem unlikely to alter transporter stability or function.

ATP Binding Cassette Transporter, Subfamily G, Mem↗

Large-scale collection and characterization of promoters of human and mouse genes.

We report the generation and initial characterization of a large-scale collection of sequences of putative promoter regions (PPRs) of human and mouse genes. Based on our unique collection of 400,225 and 580,209 human and mouse full-length cDNAs, we determined exact transcriptional start sites (TSSs). Using positional information of the TSSs, we could retrieve adjacent sequences as PPRs for 8,793 and 6,875 human and mouse genes, respectively. The positions of the PPRs were 4 kb upstream to previously reported 5'-ends of cDNAs on average, demonstrating that full-length cDNA information is indispensable for this purpose. Among those PPRs supported by experimentally validated TSSs, 3,324 could be paired as mutually homologous genes between human and mouse and were used for the comprehensive comparative studies. The sequence identities in the proximal regions of the TSSs were 45% on average, and 22,794 putative transcription factor binding sites that are conserved between human and mouse were identified. The data resource created in the present work and the results of the sequences' initial characterization should lay the firm foundation for deciphering the transcriptional modulations of human genes. All the data were deposited and made available through a database for comparative studies, DBTSS.

Animals↗

Use of mass spectrometric molecular weight information to identify proteins in sequence databases.

During the last decade new ionization techniques have made it possible to measure the molecular weight of many intact proteins by mass spectrometry, and they have made it much easier to obtain a mass spectrometric peptide map of a protein. At the same time advances in protein and DNA sequencing technology are resulting in an exponential increase in the number of sequences deposited in databases. Here we investigate the possibility to use mass spectrometric data to identify proteins in databases. Searching a database by total molecular weight is found to be an easy and sometimes sufficient approach. For more specificity and for error tolerance in both the mass spectrometric data and the database information we search by partial mass spectrometric peptide map of the protein. In general, just four to six proteolytic peptides measured with a mass accuracy between 0.1 and 0.01% allow a useful search of databases such as the Protein Identification Resource (PIR). As the size of DNA and protein sequence databases grows, protein identification by partial mass spectrometric peptide maps should become increasingly powerful and may become a general method to identify and characterize proteins.

Amino Acid Sequence↗

CAGE Basic/Analysis Databases: the CAGE resource for comprehensive promoter analysis.

Cap-analysis gene expression (CAGE) Basic and Analysis Databases store an original resource produced by CAGE, which measures expression levels of transcription starting sites by sequencing large amounts of transcript 5' ends, termed CAGE tags. Millions of human and mouse high-quality CAGE tags derived from different conditions in >20 tissues consisting of >250 RNA samples are essential for identification of novel promoters and promoter characterization in the aspect of expression profile. CAGE Basic Database is a primary database of the CAGE resource, RNA samples, CAGE libraries, CAGE clone and tag sequences and so on. CAGE Analysis Database stores promoter related information, such as counts of related transcripts, CpG islands and conserved genome region. It also provides expression profiles at base pair and promoter levels. Both databases are based on the same framework, CAGE tag starting sites, tag clusters for defining promoters and transcriptional units (TUs). Their associations and TU attributes are available to find promoters of interest. These databases were provided for Functional Annotation Of Mouse 3 (FANTOM3), an international collaboration research project focusing on expanding the transcriptome and subsequent analyses. Now access is free for all users through the World Wide Web at http://fantom3.gsc.riken.jp/.

Animals↗