Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequencing Resource”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Development of a porcine brain cDNA library, EST database, and microarray resource.

Recent developments in expressed sequence tag (EST) and cDNA microarray technology have had a dramatic impact on the ability of scientists to study responses of thousands of genes to internal and external stimuli. In neurobiology, studies of the human brain have been expanding rapidly by use of functional genomics techniques. To enhance these studies and allow use of a porcine brain model, a normalized porcine brain cDNA library (PBL) has been generated and used as a base for EST discovery and microarray generation. In this report, we discuss initial sequence analysis of 965 clones from this resource. Our data revealed that library normalization successfully reduced the number of clones representing highly abundant cDNA species and overall clone redundancy. Cluster analysis revealed over 800 unique cDNA species representing a redundancy rate for the normalized library of 6.9% compared with 29.4% before normalization. Sequence information, BLAST results, and TIGR cluster matches for these ESTs are publicly available via a web-accessible database (http://nbfgc.msu.edu). A cDNA microarray was created using 877 unique porcine brain EST amplicons spotted in triplicate on glass slides. This microarray was assessed by performing a series of experiments designed to test hybridization efficiency and false-positive rate. Our results indicate that the PBL cDNA microarray is a robust tool for studies of brain gene expression using swine as a model system.

Animals↗

OWL--a non-redundant composite protein sequence database.

A comprehensive, non-redundant composite protein sequence database is described. The database, OWL, is an amalgam of data from six publicly-available primary sources, and is generated using strict redundancy criteria. The database is updated monthly and its size has increased almost eight-fold in the last six years: the current version contains > 76,000 entries. For added flexibility, OWL is distributed with a tailor-made query language, together with a number of programs for database exploration, information retrieval and sequence analysis, which together form an integrated database and software resource for protein sequences.

Amino Acid Sequence↗

[Genetic diversity of Dactylis glomerata germplasm resources detected by Inter-simple Sequence Repeats (ISSRS) molecular markers].

Inter-simple Sequence Repeat (ISSR) molecular markers were used to detect the genetic diversity among 50 materials of Dactylis glomerata collected from China and other countries. Twelve primers produced 101 polymorphic bands, averaged 8.41 bands each primer pair. The average percentage of polymorpgic bands was 86.3.8%, and the range of GS (define) was 0.6116-0.9290, indicating a rich genetic diversity of D. glomerata. Based on the cluster and principal component analyses on the genetic characteristics, D. glomerata could be divided into 5 groups according to the nearest phylogenetic relationship. In most cases, accessions from the same continent were classified into the same group, the accessions from China and the United States belong to the different groups, respectively, indicating the geographical distribution of genetic diversity of D. glomerata. The present paper also discussed collection and conservation of germplasm resources in D. glomerata.

Conservation of Natural Resources↗

The internal transcribed spacer 2 database--a web server for (not only) low level phylogenetic analyses.

The internal transcribed spacer 2 (ITS2) is a phylogenetic marker which has been of broad use in generic and infrageneric level classifications, as its sequence evolves comparably fast. Only recently, it became clear, that the ITS2 might be useful even for higher level systematic analyses. As the secondary structure is highly conserved within all eukaryotes it serves as a valuable template for the construction of highly reliable sequence-structure alignments, which build a fundament for subsequent analyses. Thus, any phylogenetic study using ITS2 has to consider both sequence and structure. We have integrated a homology based RNA structure prediction algorithm into a web server, which allows the detection and secondary structure prediction for ITS2 in any given sequence. Furthermore, the resource contains more than 25,000 pre-calculated secondary structures for the currently known ITS2 sequences. These can be taxonomically searched and browsed. Thus, our resource could become a starting point for ITS2-based phylogenetic analyses and is therefore complementary to databases of other phylogenetic markers, which focus on higher level analyses. The current version of the ITS2 database can be accessed via http://its2.bioapps.biozentrum.uni-wuerzburg.de.

DNA, Ribosomal Spacer↗

ABCdb: an online resource for ABC transporter repertories from sequenced archaeal and bacterial genomes.

The ATP-binding cassette (ABC) transporters are one of the major classes of active transporters. They are widespread in archaea, bacteria, and eukaryota, indicating that they have arisen early in evolution. They are involved in many essential physiological processes, but the majority import or export a wide variety of compounds across cellular membranes. These systems share a common architecture composed of four (exporters) or five (importers) domains. To identify and reconstruct functional ABC transporters encoded by archaeal and bacterial genomes, we have developed a bioinformatic strategy. Cross-reference to the transport classification system is used to predict the type of compound transported. A high quality of annotation is achieved by manual verification of the predictions. However, in order to face the rapid increase in the number of published genomes, we also include analyses of genomes issuing directly from the automated strategy. Querying the database (http://www-abcdb.biotoul.fr) allows to easily retrieve ABC transporter repertories and related data. Additional query tools have been developed for the analysis of the ABC family from both functional and evolutionary perspectives.

ATP-Binding Cassette Transporters↗

ELXR: a resource for rapid exon-directed sequence analysis.

ELXR (Exon Locator and Extractor for Resequencing) streamlines the process of determining exon/intron boundaries and designing PCR and sequencing primers for high-throughput resequencing of exons. We have pre-computed ELXR primer sets for all exons identified from the human, mouse, and rat mRNA reference sequence (RefSeq) public databases curated by the National Center for Biotechnology Information. The resulting exon-flanking PCR primer pairs have been compiled into a system called ELXRdb, which may be searched by keyword, gene name or RefSeq accession number.

Algorithms↗

Optical mapping of Plasmodium falciparum chromosome 2.

Detailed restriction maps of microbial genomes are a valuable resource in genome sequencing studies but are toilsome to construct by contig construction of maps derived from cloned DNA. Analysis of genomic DNA enables large stretches of the genome to be mapped and circumvents library construction and associated cloning artifacts. We used pulsed-field gel electrophoresis purified Plasmodium falciparum chromosome 2 DNA as the starting material for optical mapping, a system for making ordered restriction maps from ensembles of individual DNA molecules. DNA molecules were bound to derivatized glass surfaces, cleaved with NheI or BamHI, and imaged by digital fluorescence microscopy. Large pieces of the chromosome containing ordered DNA restriction fragments were mapped. Maps were assembled from 50 molecules producing an average contig depth of 15 molecules and high-resolution restriction maps covering the entire chromosome. Chromosome 2 was found to be 976 kb by optical mapping with NheI, and 946 kb with BamHI, which compares closely to the published size of 947 kb from large-scale sequencing. The maps were used to further verify assemblies from the plasmid library used for sequencing. Maps generated in silico from the sequence data were compared to the optical mapping data, and good correspondence was found. Such high-resolution restriction maps may become an indispensable resource for large-scale genome sequencing projects.

Animals↗

A Drosophila full-length cDNA resource.

BACKGROUND: A collection of sequenced full-length cDNAs is an important resource both for functional genomics studies and for the determination of the intron-exon structure of genes. Providing this resource to the Drosophila melanogaster research community has been a long-term goal of the Berkeley Drosophila Genome Project. We have previously described the Drosophila Gene Collection (DGC), a set of putative full-length cDNAs that was produced by generating and analyzing over 250,000 expressed sequence tags (ESTs) derived from a variety of tissues and developmental stages. RESULTS: We have generated high-quality full-insert sequence for 8,921 clones in the DGC. We compared the sequence of these clones to the annotated Release 3 genomic sequence, and identified more than 5,300 cDNAs that contain a complete and accurate protein-coding sequence. This corresponds to at least one splice form for 40% of the predicted D. melanogaster genes. We also identified potential new cases of RNA editing. CONCLUSIONS: We show that comparison of cDNA sequences to a high-quality annotated genomic sequence is an effective approach to identifying and eliminating defective clones from a cDNA collection and ensure its utility for experimentation. Clones were eliminated either because they carry single nucleotide discrepancies, which most probably result from reverse transcriptase errors, or because they are truncated and contain only part of the protein-coding sequence.

Amino Acid Sequence↗

Structure-guided recombination creates an artificial family of cytochromes P450.

Creating artificial protein families affords new opportunities to explore the determinants of structure and biological function free from many of the constraints of natural selection. We have created an artificial family comprising 3,000 P450 heme proteins that correctly fold and incorporate a heme cofactor by recombining three cytochromes P450 at seven crossover locations chosen to minimize structural disruption. Members of this protein family differ from any known sequence at an average of 72 and by as many as 109 amino acids. Most (>73%) of the properly folded chimeric P450 heme proteins are catalytically active peroxygenases; some are more thermostable than the parent proteins. A multiple sequence alignment of 955 chimeras, including both folded and not, is a valuable resource for sequence-structure-function studies. Logistic regression analysis of the multiple sequence alignment identifies key structural contributions to cytochrome P450 heme incorporation and peroxygenase activity and suggests possible structural differences between parents CYP102A1 and CYP102A2.

Amino Acid Sequence↗

The WiscDsLox T-DNA collection: an arabidopsis community resource generated by using an improved high-throughput T-DNA sequencing pipeline.

We have developed a new community resource, called the WiscDsLox collection, for performing reverse-genetic analysis in arabidopsis. This resource is composed of 10,459 T-DNA lines generated using the Arabidopsis thaliana ecotype Columbia. The flanking sequence tag for each T-DNA insertion has been deposited in public databases, and seed for each line is currently available from the Arabidopsis Biological Resource Center. The pDsLox vector used to create this new population contains a Ds transposon and Cre/Lox recombination sites. Each WiscDsLox line therefore has the potential to serve as a launch-pad for performing local saturation mutagenesis by mobilization of the Ds element. In addition, Cre-Lox recombination between the T-DNA and a transposed Ds element should enable targeted deletion of specific genomic regions. We generated the WiscDsLox collection using an improved high-throughput pipeline that streamlines analysis of large numbers of independent Arabidopsis thaliana (L.) Hyenh. lines. In this paper we describe the details of this novel method and also provide potential users of WiscDsLox T-DNA lines with useful background information about this collection. Experiments to characterize the utility of the Ds transposon and Cre/Lox elements present in the WiscDsLox lines are in progress and will be reported in the future.

Arabidopsis↗

The PANTHER database of protein families, subfamilies, functions and pathways.

PANTHER is a large collection of protein families that have been subdivided into functionally related subfamilies, using human expertise. These subfamilies model the divergence of specific functions within protein families, allowing more accurate association with function (ontology terms and pathways), as well as inference of amino acids important for functional specificity. Hidden Markov models (HMMs) are built for each family and subfamily for classifying additional protein sequences. The latest version, 5.0, contains 6683 protein families, divided into 31,705 subfamilies, covering approximately 90% of mammalian protein-coding genes. PANTHER 5.0 includes a number of significant improvements over previous versions, most notably (i) representation of pathways (primarily signaling pathways) and association with subfamilies and individual protein sequences; (ii) an improved methodology for defining the PANTHER families and subfamilies, and for building the HMMs; (iii) resources for scoring sequences against PANTHER HMMs both over the web and locally; and (iv) a number of new web resources to facilitate analysis of large gene lists, including data generated from high-throughput expression experiments. Efforts are underway to add PANTHER to the InterPro suite of databases, and to make PANTHER consistent with the PIRSF database. PANTHER is now publicly available without restriction at http://panther.appliedbiosystems.com.

Animals↗

The Biobank Rare Variant consortium powers the discovery of rare genetic associations through global collaboration.

Rare coding variants can have large effects on disease risk and provide direct routes from human genetics to disease mechanisms and therapeutic targets, but their discovery is constrained by sample size, particularly for low-prevalence diseases. Here we establish the Biobank Rare Variant Analysis (BRaVa) consortium, a global rare variant association resource that integrates sequencing and linked health-record data from ten biobanks and cohorts comprising over 1.2 million individuals across diverse ancestries. We performed gene-based meta-analyses of rare coding variation across 33 clinical endpoints and 11 quantitative traits. Aggregating evidence across biobanks and ancestries identified 514 gene-trait associations, including 31 not previously reported in prior studies or curated association resources following systematic literature review. Notably, 36.1% of gene-level associations were undetectable in any individual biobank, and 91 emerged only through cross-ancestry meta-analysis, demonstrating that federated integration enables discovery beyond the reach of single cohorts. Similar gains were observed at the variant level, where 25.0% of phenotype-locus associations were detectable only through meta-analysis. Effect size estimates were correlated across ancestries with concordant directions of effect, supporting the generalizability of rare variant associations. The identified signals implicate pathways involved in transcriptional and epigenetic regulation, metabolism, vascular and epithelial biology, and immune function, highlighting rare coding variation as an engine for biological discovery across medical record phenotypes. For example, damaging variation in ANKRD12 implicates inflammatory transcriptional dysregulation in asthma and chronic obstructive pulmonary disease, and ultra-rare predicted loss-of-function variants in NAA15 link protein acetylation processes to type 2 diabetes risk. BRaVa establishes a scalable framework and freely available community resource for rare variant meta-analysis across global biobanks. Public release of gene- and variant-level association summary statistics provides a reference map of rare coding variant associations to support disease gene discovery, biological interpretation, and therapeutic target prioritization as sequencing-linked health-record resources continue to expand.

Journal Article↗

The TIGR Gene Indices: analysis of gene transcript sequences in highly sampled eukaryotic species.

While genome sequencing projects are advancing rapidly, EST sequencing and analysis remains a primary research tool for the identification and categorization of gene sequences in a wide variety of species and an important resource for annotation of genomic sequence. The TIGR Gene Indices (http://www.tigr.org/tdb/tgi. shtml) are a collection of species-specific databases that use a highly refined protocol to analyze EST sequences in an attempt to identify the genes represented by that data and to provide additional information regarding those genes. Gene Indices are constructed by first clustering, then assembling EST and annotated gene sequences from GenBank for the targeted species. This process produces a set of unique, high-fidelity virtual transcripts, or Tentative Consensus (TC) sequences. The TC sequences can be used to provide putative genes with functional annotation, to link the transcripts to mapping and genomic sequence data, to provide links between orthologous and paralogous genes and as a resource for comparative sequence analysis.

Animals↗

Analysis of VH gene sequences using two web-based immunogenetics resources gives different results, but the affinity maturation status of chronic lymphocytic leukaemia clones as assessed from either of the resulting data sets has no prognostic significance.

Some cellular and molecular features of chronic lymphocytic leukaemia (CLL) cells that are associated with prognosis may reflect the context within which their progenitors encountered antigen. It follows that the nature of antigen drive in CLL could influence the clinical course and we were prompted to assess the impact, if any, of affinity maturation (an antigen-driven process) on prognosis. Statistical models for assessing affinity maturation status are typically applied to V(H) gene sequence data analysed using a web-based resource like IMGT or VBASE. Since these resources differ with respect to some key relevant features, we evaluated a cohort of CLL cases by applying statistical models to V(H) data derived from both IMGT and VBASE. Important differences between the resulting data sets became apparent. These resulted from database variance and because IMGT and VBASE define complementarity-determining and framework regions (CDRs, FRs) in different ways. Thus, the numbers of mutations identified and their distribution between CDRs/FRs varied between the data sets for the majority of clones. Consequently, two different but overlapping sets of cases with evidence of affinity maturation were defined. Notwithstanding their differences, no significant associations of affinity maturation status with CD38 expression, p53 functional status or survival were identifiable in either data set.

ADP-ribosyl Cyclase↗

Integrating computationally assembled mouse transcript sequences with the Mouse Genome Informatics (MGI) database.

Databases of experimentally generated and computationally derived transcript sequences are valuable resources for genome analysis and annotation. The utility of such databases is enhanced when the sequences they contain are integrated with such biological information as genomic location, gene function, gene expression and phenotypic variation. We present the analysis and results of a semi-automated process of connecting transcript assemblies with highly curated biological information for mouse genes that is available through the Mouse Genome Informatics (MGI) database.

Animals↗

Wheat EST sequence assembly facilitates comparison of gene contents among plant species and discovery of novel genes.

Using a strategy requiring only modest computational resources, wheat expressed sequence tag (EST) sequences from various sources were assembled into contigs and compared with a nonredundant barley sequence assembly, with ESTs, with complete draft genome sequences of rice and Arabidopsis thaliana, and with ESTs from other plant species. These comparisons indicate that (i) wheat sequences available from public sources represent a substantial proportion of the diversity of wheat coding sequences, (ii) prediction of open reading frames in the whole genome sequence improves when supplemented with EST information from other species, (iii) a substantial number of candidates for novel genes that are unique to wheat or related species can be identified, and (iv) a smaller number of genes can be identified that are common to monocots and dicots but absent from Arabidopsis. The sequences in the last group may have been lost from Arabidopsis after descendance from a common ancestor. Examples of potential novel wheat genes and Triticeae-specific genes are presented.

Arabidopsis↗

NCBI's LocusLink and RefSeq.

The NCBI has introduced two new web resources-LocusLink and RefSeq-that facilitate retrieval of gene-based information and provide reference sequence standards. These resources are designed to provide a non-redundant view of current knowledge about human genes, transcripts and proteins. Additional information about these resources is available on the LocusLink web site at http://www.ncbi.nlm.nih.gov/LocusLink/

Database Management Systems↗

Construction of validated, non-redundant composite protein sequence databases.

A strategy has been developed for the construction of a validated, comprehensive composite protein sequence database. Entries are amalgamated from primary source data bases by a largely automated set of processes in which redundant and trivially different entries are eliminated. A modular approach has been adopted to allow scientific judgement to be used at each stage of database processing and amalgamation. Source databases are assigned a priority depending on the quality of sequence validation and commenting. Rejection of entries from the lower priority database, in each pairwise comparison of databases, is carried out according to optionally defined redundancy criteria based on sequence segment mismatches. Efficient algorithms for this methodology are embodied in the COMPO software system. COMPO has been applied for over 2 years in construction and regular updating of the OWL composite protein sequence database from the source databases NBRF-PIR, SWISS-PROT, a GenBank translation retrieved from the feature tables, NBRF-NEW, NEWAT86, PSD-KYOTO and the sequences contained in the Brookhaven protein structure databank. OWL is part of the ISIS integrated data resource of protein sequence and structure [Akrigg et al. (1988) Nature, 335, 745-746]. The modular nature of the integration process greatly facilitates the frequent updating of OWL following releases of the source databases. The extent of redundancy in these sources is revealed by the comparison process. The advantages of a robust composite database for sequence similarity searching and information retrieval are discussed.

Amino Acid Sequence↗