Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequencing Resource”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

NCBI Reference Sequence (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins.

The National Center for Biotechnology Information (NCBI) Reference Sequence (RefSeq) database (http://www.ncbi.nlm.nih.gov/RefSeq/) provides a non-redundant collection of sequences representing genomic data, transcripts and proteins. Although the goal is to provide a comprehensive dataset representing the complete sequence information for any given species, the database pragmatically includes sequence data that are currently publicly available in the archival databases. The database incorporates data from over 2400 organisms and includes over one million proteins representing significant taxonomic diversity spanning prokaryotes, eukaryotes and viruses. Nucleotide and protein sequences are explicitly linked, and the sequences are linked to other resources including the NCBI Map Viewer and Gene. Sequences are annotated to include coding regions, conserved domains, variation, references, names, database cross-references, and other features using a combined approach of collaboration and other input from the scientific community, automated annotation, propagation from GenBank and curation by NCBI staff.

Animals↗

Partial genome sequencing of Rhodococcus equi ATCC 33701.

Preliminary analysis of a partial (30% coverage) genome sequence of Rhodococcus equi has revealed a number of important features. The most notable was the extent of the homology of genes identified with those of Mycobacterium tuberculosis. The similarities in the proportion of genes devoted to fatty acid degradation and to lipid biosynthesis was a striking but not surprising finding given the relatedness of these organisms and their success as intracellular pathogens. The rapid recent improvement in understanding of virulence in M. tuberculosis and other pathogenic mycobacteria has identified a large number of genes of putative or proven importance in virulence, homologs of many of which were also identified in R. equi. Although R. equi appears to have currently unique genes, and has important differences, its similarity to M. tuberculosis supports the need to understand the basis of virulence in this organism. The partial genome sequence will be a resource for workers interested in R. equi until such time as a full genome sequence has been characterized.

Aerobiosis↗

Multilocus sequence typing method for identification and genotypic classification of pathogenic Leptospira species.

BACKGROUND: Leptospira are the parasitic bacterial organisms associated with a broad range of mammalian hosts and are responsible for severe cases of human Leptospirosis. The epidemiology of leptospirosis is complex and dynamic. Multiple serovars have been identified, each adapted to one or more animal hosts. Adaptation is a dynamic process that changes the spatial and temporal distribution of serovars and clinical manifestations in different hosts. Serotyping based on repertoire of surface antigens is an ambiguous and artificial system of classification of leptospiral agents. Molecular typing methods for the identification of pathogenic leptospires up to individual genome species level have been highly sought after since the decipherment of whole genome sequences. Only a few resources exist for microbial genotypic data based on individual techniques such as Multiple Locus Sequence Typing (MLST), but unfortunately no such databases are existent for leptospires. RESULTS: We for the first time report development of a robust MLST method for genotyping of Leptospira. Genotyping based on DNA sequence identity of 4 housekeeping genes and 2 candidate genes was analyzed in a set of 120 strains including 41 reference strains representing different geographical areas and from different sources. Of the six selected genes, adk, icdA and secY were significantly more variable whereas the LipL32 and LipL41 coding genes and the rrs2 gene were moderately variable. The phylogenetic tree clustered the isolates according to the genome-based species. CONCLUSION: The main advantages of MLST over other typing methods for leptospires include reproducibility, robustness, consistency and portability. The genetic relatedness of the leptospires can be better studied by the MLST approach and can be used for molecular epidemiological and evolutionary studies and population genetics.

Animals↗

Identification of polymorphic tandem repeats by direct comparison of genome sequence from different bacterial strains: a web-based resource.

BACKGROUND: Polymorphic tandem repeat typing is a new generic technology which has been proved to be very efficient for bacterial pathogens such as B. anthracis, M. tuberculosis, P. aeruginosa, L. pneumophila, Y. pestis. The previously developed tandem repeats database takes advantage of the release of genome sequence data for a growing number of bacteria to facilitate the identification of tandem repeats. The development of an assay then requires the evaluation of tandem repeat polymorphism on well-selected sets of isolates. In the case of major human pathogens, such as S. aureus, more than one strain is being sequenced, so that tandem repeats most likely to be polymorphic can now be selected in silico based on genome sequence comparison. RESULTS: In addition to the previously described general Tandem Repeats Database, we have developed a tool to automatically identify tandem repeats of a different length in the genome sequence of two (or more) closely related bacterial strains. Genome comparisons are pre-computed. The results of the comparisons are parsed in a database, which can be conveniently queried over the internet according to criteria of practical value, including repeat unit length, predicted size difference, etc. Comparisons are available for 16 bacterial species, and the orthopox viruses, including the variola virus and three of its close neighbors. CONCLUSIONS: We are presenting an internet-based resource to help develop and perform tandem repeats based bacterial strain typing. The tools accessible at http://minisatellites.u-psud.fr now comprise four parts. The Tandem Repeats Database enables the identification of tandem repeats across entire genomes. The Strain Comparison Page identifies tandem repeats differing between different genome sequences from the same species. The "Blast in the Tandem Repeats Database" facilitates the search for a known tandem repeat and the prediction of amplification product sizes. The "Bacterial Genotyping Page" is a service for strain identification at the subspecies level.

Bacteria↗

2058 expressed sequence tags (ESTs) from a human fetal lung cDNA library.

ESTs (expressed sequence tags) provide complementary resources for structural and functional analyses of the human genome. We have performed single-pass sequencing of 2058 randomly selected, directionally cloned cDNAs isolated from a fetal-lung cDNA library constructed with oligo(dT) primers. Computer analyses of the 5'-end sequences revealed that 60.4% of the clones were considered to be identical to previously reported human genes or ESTs; 9.0% of them showed significant homology to known genes in human, other mammals, or lower organisms; 30.6% showed no homology to any genes or DNA sequences in the public database. These data and reagents will be useful for future investigations of gene expression during prenatal development of human lung.

Amino Acid Sequence↗

UniProt archive.

UniProt Archive (UniParc) is the most comprehensive, non-redundant protein sequence database available. Its protein sequences are retrieved from predominant, publicly accessible resources. All new and updated protein sequences are collected and loaded daily into UniParc for full coverage. To avoid redundancy, each unique sequence is stored only once with a stable protein identifier, which can be used later in UniParc to identify the same protein in all source databases. When proteins are loaded into the database, database cross-references are created to link them to the origins of the sequences. As a result, performing a sequence search against UniParc is equivalent to performing the same search against all databases cross-referenced by UniParc. UniParc contains only protein sequences and database cross-references; all other information must be retrieved from the source databases.

Amino Acid Sequence↗

Characterization of open reading frame-expressed sequence tags generated from Bos indicus and B. taurus mammary gland cDNA libraries.

Sequence-based gene expression data are used to interpret results from functional genomic and proteomics studies. Although more than 300 000 bovine-expressed sequence tags (ESTs) are available in public databases, a more thorough and directed sampling of the expressed genome is needed to identify new transcripts and improve assembly and annotation of existing transcript sequences. Accordingly, we examined the utility of constructing cDNA libraries synthesized by arbitrarily primed RT-PCR of mRNA from tissues not well represented in the publicly available bovine EST database. A total of 33 cDNA libraries were constructed from healthy and infected mammary gland tissues of Brazilian Gir and Holstein cattle. This series of libraries was used to generate 6481 open reading frame-expressed sequence tags (ORESTES) that assembled into 1798 unique sequence elements of which, 1157 did not significantly match sequence assemblies available in the Bos taurus gene index. However, a total of 264 of these 1157 sequence elements aligned with mouse and human expressed sequences demonstrating that ORESTES is an effective resource for discovery of novel expressed sequences in cattle. Furthermore, comparison of the alignment position of bovine ORESTES-derived sequence elements to human gene reference sequences suggested that the priming events for cDNA synthesis more often occurred at the central portion of a transcript, which may have contributed to the relatively high rate of novel sequence discovery.

Animals↗

Accuracy of automated DNA sequencing: a multi-laboratory comparison of sequencing results.

A double-stranded (ds)DNA template of "unknown" sequence was distributed to approximately 80 core DNA sequencing laboratories by the Association of Biomolecular Resource Facilities (ABRF) for automated DNA sequence analysis. Forty-four different facilities responded with 83 usable sequence submissions. These sequences were grouped by both sequencing protocol (dye-primer or dye-terminator) and whether manually edited or not. The sequences were aligned with the known sequence, and the number of correct base calls, insertions, deletions, no-calls and miscalls were determined for each group. The dye-primer sequencing protocol provided the longest and most accurate sequence. The edited dye-primer data were > 95% accurate out to 400-450 bp, while the edited dye-terminator data could call only 300-350 bases at this accuracy. However, 75% of the laboratories in this sampling preferred the dye-terminator protocol, presumably because of its versatility and convenience. Laboratories that manually edited the automatically called data were able to obtain an additional 100 bases of good sequence when the dye-primer protocol was used. Surprisingly though, editing of dye-terminator results did not increase the amount of good sequence, although the dye-terminator protocol had a superior base-calling ability within the first 100 bases of called sequence.

Autoanalysis↗

A trainable language model with potential to modulate translation rates in non-model organisms by generating upstream untranslated region sequence libraries.

Tuning protein expression in non-model organisms is often constrained by the lack of validated genetic parts and predictive design tools. Translational tuning through the modulation of upstream untranslated regions (5'-UTRs) offers a potentially organism-agnostic route, but existing methods typically rely on mechanistic assumptions, prior knowledge that may not be available in non-model contexts, or the screening of sequence libraries. Here, we present a simple generative approach for creating synthetic 5'-UTR libraries based solely on the genomic sequence statistics of any desired organism. The method uses a sliding-window n-gram language model applied to native 5'-UTR sequences to produce novel sequences that preserve organism-specific base distributions and motifs without hard-coding specific motifs or mechanistic rules into inflexible statistical templates. We have applied this approach to the model bacterium Escherichia coli and the non-model probiotic Limosilactobacillus reuteri. Libraries of approximately 1,000 sequences were generated for each organism, from which about 100 unique sequences were experimentally tested for translation of a fluorescent reporter protein. In both organisms, the synthetic libraries yielded a broad range of translation levels from this relatively small number of tested variants. Sequences derived from an organism's own genomic statistics provided a more uniformly distributed range of translation rates in that organism than sequences derived from the other species. Correlations of individual sequence performance across the two species were weak, and thermodynamic predictions of ribosome binding strength showed very little predictive power, especially in the non-model L. reuteri. The results demonstrate that simple statistical language model approaches applied to genomic data can generate functional translational regulatory sequence libraries without detailed mechanistic knowledge or explicit reference to consensus motifs. The approach requires minimal computational resources, avoids reproducing native sequences, and can be readily applied to any organism with a sequenced genome. This strategy may lower technical barriers to expression tuning in non-model organisms.

5' Untranslated Regions↗

The complete genome sequence of a dog: a perspective.

A complete, high-quality reference sequence of a dog genome was recently produced by a team of researchers led by the Broad Institute, achieving another major milestone in deciphering the genomic landscape of mammalian organisms. The genome sequence provides an indispensable resource for comparative analysis and novel insights into dog and human evolution and history. Together with the survey sequence of a poodle previously published in 2003, the two dog genome sequences allowed identification of more than 2.5 million single nucleotide polymorphisms within and between dog breeds, which can be used in evolutionary analysis, behavioral studies and disease gene mapping.(1)

Animals↗

Comparative genomics for the investigation of autoimmune diseases.

The complete DNA sequence of the human genome and of several related mammals are now available, due to the investments of enormous resources and advances in sequencing technology. Novel technologies have been developed to compare multiple genomes with each other, thus specifying regions of sequence similarity among mammals and with their pathogens. Larger blocks of sequence similarity (syntenic regions) have been determined and made publicly available. In many ways, novel insights can be gained by such data when combining external genetic or clinical information for these syntenic loci. These novel tools have proven to be successful in inferring functional equivalence between loci of multiple genomes. This review reports on the role of comparative genomics in research on autoimmune diseases, a field with strong dependencies on animal models of human diseases and the problem of an adequate information transfer between multiple organisms and research areas.

Animals↗

Sequence variation database project at the European Bioinformatics Institute.

The sequence variation project at EBI aims to create a unified resource for browsing and searching sequence differences. Technical advances in reading in new data types and in validating and cross-referencing entries are reported. It is suggested that the hardest problems in unifying mutation databases are related to intellectual property rights. The concept of copylefting is introduced as a potential solution to these.

Computational Biology↗

A sequence-ready BAC clone contig of human chromosome 10p15 spanning the loss of heterozygosity region in glioma.

Deletion of chromosome 10 is one of the most common chromosomal alterations in glioma. At 10p15, the telomeric region of the short arm of chromosome 10, loss of heterozygosity (LOH) has been frequently observed by microsatellite analysis, suggesting the presence of a tumor suppressor gene. We examined LOH in 34 gliomas on chromosome 10, and frequent LOH on 10p was detected on 10p15, in agreement with deletion mapping studies on chromosome 10. We then constructed a bacterial artificial chromosome (BAC) clone contig covering the critical region, which spanned the interval between D10S249 and D10S533 on 10p15. The map contained 68 BAC clones connected by 74 sequenced tag sites (STSs) and covered approximately 2.7 Mb, with one gap. A total of 74 STSs, including 6 microsatellite markers, 29 expressed sequenced tags (ESTs), and 39 BAC end STSs, were physically arranged. Twenty-eight ESTs were mapped in the interval between D10S249 and D10S559 (approximately 1200 kb), and another EST was mapped in the interval between D10S559 and D10S533 (approximately 1300 kb). This sequence-ready BAC clone contig map will be a basic resource for high-quality sequencing and positional cloning of the putative tumor suppressor gene at 10p15 in glioma.

Base Sequence↗

Sequence learning under dual-task conditions: alternatives to a resource-based account.

In two experiments with the serial reaction-time task, participants were presented with deterministic or probabilistic sequences under single- or dual-task conditions. Experiment 1 showed that learning of a probabilistic structure was not impaired over a first session by performing a counting task, but that such an interference arose over a second session, when the knowledge was tested under single-task conditions. In contrast, the effects of the secondary task arose earlier for participants exposed to deterministic sequences. This difference between deterministic and probabilistic sequences disappeared in Experiment 2, where the counting task was performed on tones associated to the locations. Comparisons between sessions indicated that the secondary task affected not only the expression but also the acquisition of sequence learning, and that greater interference was observed in those conditions that yielded more explicit knowledge. These results suggest that the effects of a dual task on the measures of implicit sequence learning may be partly due to the intrusion of explicit knowledge and partly due to the disruption of the sequence produced by the inclusion of random events.

Cues↗

PIR-ALN: a database of protein sequence alignments.

MOTIVATION: The Protein Information Resource (PIR) maintains a database of annotated and curated alignments in order to visually represent interrelationships among sequences in the PIR-International Protein Sequence Database, to spread and standardize protein names, features and keywords among members of a family or superfamily, and to aid us in classifying sequences, in identifying conserved regions, and in defining new homology domains. RESULTS: Release 22.0, (December 1998), of the PIR-ALN database contains a total of 3806 alignments, including 1303 superfamily, 2131 family and 372 homology domain alignments. This is an appropriate dataset to develop and extract patterns, test profiles, train neural networks or build Hidden Markov Models (HMMs). These alignments can be used to standardize and spread annotation to newer members by homology, as well as to understand the modular architecture of multidomain proteins. PIR-ALN includes 529 alignments that can be used to develop patterns not represented in PROSITE, Blocks, PRINTS and Pfam databases. The ATLAS information retrieval system can be used to browse and query the PIR-ALN alignments. AVAILABILITY: PIR-ALN is currently being distributed as a single ASCII text file along with the title, member, species, superfamily and keyword indexes. The quarterly and weekly updates can be accessed via the WWW at pir.georgetown.edu. The quarterly updates can also be obtained by anonymous FTP from the PIR FTP site at NBRF.Georgetown.edu, directory [ANONYMOUS.PIR.ALIGNMENT].

Amino Acid Sequence↗

Genomics of hybrid poplar (Populus trichocarpax deltoides) interacting with forest tent caterpillars (Malacosoma disstria): normalized and full-length cDNA libraries, expressed sequence tags, and a cDNA microarray for the study of insect-induced defences in poplar.

As part of a genomics strategy to characterize inducible defences against insect herbivory in poplar, we developed a comprehensive suite of functional genomics resources including cDNA libraries, expressed sequence tags (ESTs) and a cDNA microarray platform. These resources are designed to complement the existing poplar genome sequence and poplar (Populus spp.) ESTs by focusing on herbivore- and elicitor-treated tissues and incorporating normalization methods to capture rare transcripts. From a set of 15 standard, normalized or full-length cDNA libraries, we generated 139,007 3'- or 5'-end sequenced ESTs, representing more than one-third of the c. 385,000 publicly available Populus ESTs. Clustering and assembly of 107,519 3'-end ESTs resulted in 14,451 contigs and 20,560 singletons, altogether representing 35,011 putative unique transcripts, or potentially more than three-quarters of the predicted c. 45,000 genes in the poplar genome. Using this EST resource, we developed a cDNA microarray containing 15,496 unique genes, which was utilized to monitor gene expression in poplar leaves in response to herbivory by forest tent caterpillars (Malacosoma disstria). After 24 h of feeding, 1191 genes were classified as up-regulated, compared to only 537 down-regulated. Functional classification of this induced gene set revealed genes with roles in plant defence (e.g. endochitinases, Kunitz protease inhibitors), octadecanoid and ethylene signalling (e.g. lipoxygenase, allene oxide synthase, 1-aminocyclopropane-1-carboxylate oxidase), transport (e.g. ABC proteins, calreticulin), secondary metabolism [e.g. polyphenol oxidase, isoflavone reductase, (-)-germacrene D synthase] and transcriptional regulation [e.g. leucine-rich repeat transmembrane kinase, several transcription factor classes (zinc finger C3H type, AP2/EREBP, WRKY, bHLH)]. This study provides the first genome-scale approach to characterize insect-induced defences in a woody perennial providing a solid platform for functional investigation of plant-insect interactions in poplar.

Animals↗

VirGen: a comprehensive viral genome resource.

VirGen is a comprehensive viral genome resource that organizes the 'sequence space' of viral genomes in a structured fashion. It has been developed with the objective of serving as an annotated and curated database comprising complete genome sequences of viruses, value-added derived data and data mining tools. The current release (v1.1) contains 559 complete genomes in addition to 287 putative genomes of viruses belonging to eight viral families for which the host range includes animals and plants. Viral genomes in VirGen are annotated using sequence-based Bioinformatics approaches. The genomic data is also curated to identify 'alternate names' of viral proteins, where available. VirGen archives the results of comparisons of genomes, proteomes and individual proteins within and between viral species. It is the first resource to provide phylogenetic trees of viral species computed using whole-genome sequence data. The module of predicted B-cell antigenic determinants in VirGen is an attempt to link the genome to its vaccinome. Comparative genome analysis data facilitate the study of genome organization and evolution of viruses, which would have implications in applied research to identify candidates for the design of vaccines and antiviral drugs. VirGen is a relational database and is available at http://bioinfo. ernet.in/virgen/virgen.html.

Antigens, Viral↗

Linking proteome and genome: how to identify parasite proteins.

Parasite genome projects are generating an avalanche of sequence data. If this resource is to be exploited effectively for drug and vaccine design, there is an urgent need to make the link between these DNA sequences and the functional proteins of the parasite, which they encode. Here, we seek to demystify the revolutionary advances in protein identification based on mass spectrometry.

Amino Acid Sequence↗