Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequencing Resource”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Draft versus finished sequence data for DNA and protein diagnostic signature development.

Sequencing pathogen genomes is costly, demanding careful allocation of limited sequencing resources. We built a computational Sequencing Analysis Pipeline (SAP) to guide decisions regarding the amount of genomic sequencing necessary to develop high-quality diagnostic DNA and protein signatures. SAP uses simulations to estimate the number of target genomes and close phylogenetic relatives (near neighbors or NNs) to sequence. We use SAP to assess whether draft data are sufficient or finished sequencing is required using Marburg and variola virus sequences. Simulations indicate that intermediate to high-quality draft with error rates of 10(-3)-10(-5) (approximately 8x coverage) of target organisms is suitable for DNA signature prediction. Low-quality draft with error rates of approximately 1% (3x to 6x coverage) of target isolates is inadequate for DNA signature prediction, although low-quality draft of NNs is sufficient, as long as the target genomes are of high quality. For protein signature prediction, sequencing errors in target genomes substantially reduce the detection of amino acid sequence conservation, even if the draft is of high quality. In summary, high-quality draft of target and low-quality draft of NNs appears to be a cost-effective investment for DNA signature prediction, but may lead to underestimation of predicted protein signatures.

Computational Biology↗

Fifty-four new gene-based canine microsatellite markers.

Fifty-four new markers were developed to fill in gaps in the current map of canine microsatellites and to complement existing markers that may not be sufficiently informative in highly inbred canine pedigrees. Canine genes contained on the radiation hybrid map were used to obtain the sequence of the human homolog. A BLAST search versus the canine whole genome shotgun (wgs) sequence resource was used to obtain the sequence of the canine genomic contigs containing the homolog of the corresponding human gene. Canine sequences that contained microsatellites and mapped back to the correct location in the human genome were used to design primers for amplification of the microsatellites from canine genomic DNA. Heterozygosities of the markers were tested by genotyping grandparental DNAs obtained from the Nestle Purina Reference family DNA distribution center plus DNAs from unrelated Bouviers and Irish wolfhounds. Canine map positions of markers on the July 2004 freeze of the canine genome assembly were determined by in silico PCR or BLAST.

Animals↗

The EMBL Nucleotide Sequence Database.

The EMBL Nucleotide Sequence Database (http://www.ebi.ac.uk/embl.html) constitutes Europe's primary nucleotide sequence resource. Main sources for DNA and RNA sequences are direct submissions from individual researchers, genome sequencing projects and patent applications. While automatic procedures allow incorporation of sequence data from large-scale genome sequencing centres and from the European Patent Office (EPO), the preferred submission tool for individual submitters is Webin (WWW). Through all stages, dataflow is monitored by EBI biologists communicating with the sequencing groups. In collaboration with DDBJ and GenBank the database is produced, maintained and distributed at the European Bioinformatics Institute (EBI). Database releases are produced quarterly and are distributed on CD-ROM. Network services allow access to the most up-to-date data collection via Internet and World Wide Web interface. EBI's Sequence Retrieval System (SRS) is a Network Browser for Databanks in Molecular Biology, integrating and linking the main nucleotide and protein databases, plus many specialised databases. For sequence similarity searching a variety of tools (e.g. Blitz, Fasta, Blast etc) are available for external users to compare their own sequences against the most currently available data in the EMBL Nucleotide Sequence Database and SWISS-PROT.

Amino Acid Sequence↗

A gene-based high-resolution comparative radiation hybrid map as a framework for genome sequence assembly of a bovine chromosome 6 region associated with QTL for growth, body composition, and milk performance traits.

BACKGROUND: A number of different quantitative trait loci (QTL) for various phenotypic traits, including milk production, functional, and conformation traits in dairy cattle as well as growth and body composition traits in meat cattle, have been mapped consistently in the middle region of bovine chromosome 6 (BTA6). Dense genetic and physical maps and, ultimately, a fully annotated genome sequence as well as their mutual connections are required to efficiently identify genes and gene variants responsible for genetic variation of phenotypic traits. A comprehensive high-resolution gene-rich map linking densely spaced bovine markers and genes to the annotated human genome sequence is required as a framework to facilitate this approach for the region on BTA6 carrying the QTL. RESULTS: Therefore, we constructed a high-resolution radiation hybrid (RH) map for the QTL containing chromosomal region of BTA6. This new RH map with a total of 234 loci including 115 genes and ESTs displays a substantial increase in loci density compared to existing physical BTA6 maps. Screening the available bovine genome sequence resources, a total of 73 loci could be assigned to sequence contigs, which were already identified as specific for BTA6. For 43 loci, corresponding sequence contigs, which were not yet placed on the bovine genome assembly, were identified. In addition, the improved potential of this high-resolution RH map for BTA6 with respect to comparative mapping was demonstrated. Mapping a large number of genes on BTA6 and cross-referencing them with map locations in corresponding syntenic multi-species chromosome segments (human, mouse, rat, dog, chicken) achieved a refined accurate alignment of conserved segments and evolutionary breakpoints across the species included. CONCLUSION: The gene-anchored high-resolution RH map (1 locus/300 kb) for the targeted region of BTA6 presented here will provide a valuable platform to guide high-quality assembling and annotation of the currently existing bovine genome sequence draft to establish the final architecture of BTA6. Hence, a sequence-based map will provide a key resource to facilitate prospective continued efforts for the selection and validation of relevant positional and functional candidates underlying QTL for milk production and growth-related traits mapped on BTA6 and on similar chromosomal regions from evolutionary closely related species like sheep and goat. Furthermore, the high-resolution sequence-referenced BTA6 map will enable precise identification of multi-species conserved chromosome segments and evolutionary breakpoints in mammalian phylogenetic studies.

Animals↗

CpG island libraries from human chromosomes 18 and 22: landmarks for novel genes.

CpG islands are found at the 5' end of approximately 60% of human genes and so are important genomic landmarks. They are concentrated in early-replicating, highly acetylated gene-rich regions. With respect to CpG island content, human Chrs 18 and 22 are very different from each other: Chr 18 appears to be CpG island poor, whereas Chr 22 appears to be CpG island rich. We have constructed and validated CpG island libraries from flow-sorted Chrs 18 and 22 and used these to estimate the difference in number of CpG islands found on these two chromosomes. These libraries contain normalized collections of sequences from the 5' end of genes. Clones from the libraries were sequenced and compared with the sequence databases; one third matched ESTs, thus anchoring these ESTs at the 5' end of their gene. However, it was striking that many clones either had no match or matched only existing CpG island clones. This suggests that a significant proportion of 5' gene sequences are absent from databases, presumably either because they are difficult to clone or the gene is poorly expressed and/or has a restricted expression pattern. This point should be taken into consideration if the currently available libraries are those used for the elucidation of complete, as opposed to partial, gene sequences. The Chr 18 and 22 CpG island libraries are a sequence resource for the isolation of such 5' gene sequences from specific human chromosomes.

Base Sequence↗

PANAL: an integrated resource for Protein sequence ANALysis.

SUMMARY: We present PANAL, an integrated resource for protein sequence analysis. The tool allows the user to simultaneously search a protein sequence for motifs from several databases, and to view the result as an intuitive graphical summary.

Computational Biology↗

Exploitation of pepper EST-SSRs and an SSR-based linkage map.

As genome and cDNA sequencing projects progress, a tremendous amount of sequence information is becoming publicly available. These sequence resources can be exploited for gene discovery and marker development. Simple sequence repeat (SSR) markers are among the most useful because of their great variability, abundance, and ease of analysis. By in silico analysis of 10,232 non-redundant expressed sequence tags (ESTs) in pepper as a source of SSR markers, 1,201 SSRs were found, corresponding to one SSR in every 3.8 kb of the ESTs. Eighteen percent of the SSR-ESTs were dinucleotide repeats, 66.0% were trinucleotide, 7.7% tetranucleotide, and 8.2% pentanucleotide; AAG (14%) and AG (12.4%) motifs were the most abundant repeat types. Based on the flanking sequences of these 1,201 SSRs, 812 primer pairs that satisfied melting temperature conditions and PCR product sizes were designed. 513 SSRs (63.1%) were successfully amplified and 150 of them (29.2%) showed polymorphism between Capsicum annuum 'TF68' and C. chinense 'Habanero'. Dinucleotide SSRs and EST-SSR markers containing AC-motifs were the most polymorphic. Polymorphism increased with repeat length and repeat number. The polymorphic EST-SSRs were mapped onto the previously generated pepper linkage map, using 107 F(2) individuals from an interspecific cross of TF68 x Habanero. One-hundred and thirtynine EST-SSRs were located on the linkage map in addition to 41 previous SSRs and 63 RFLP markers, forming 14 linkage groups (LGs) and spanning 2,201.5 cM. The EST-SSR markers were distributed over all the LGs. This SSR-based map will be useful as a reference map in Capsicum and should facilitate the use of molecular markers in pepper breeding.

Capsicum↗

Systematic sequencing of cDNA clones using the transposon Tn5.

In parallel with the production of genomic sequence data, attention is being focused on the generation of comprehensive cDNA-sequence resources. Such efforts are increasingly emphasizing the production of high-accuracy sequence corresponding to the entire insert of cDNA clones, especially those presumed to reflect the full-length mRNA. The complete sequencing of cDNA clones on a large scale presents unique challenges because of the generally small, yet heterogeneous, sizes of the cloned inserts. We have developed a strategy for high-throughput sequencing of cDNA clones using the transposon Tn5. This approach has been tailored for implementation within an existing large-scale 'shotgun-style' sequencing program, although it could be readily adapted for use in virtually any sequencing environment. In addition, we have developed a modified version of our strategy that can be applied to cDNA clones with large cloning vectors, thereby overcoming a potential limitation of transposon-based approaches. Here we describe the details of our cDNA-sequencing pipeline, including a summary of the experience in sequencing more than 4200 cDNA clones to produce more than 8 million base pairs of high-accuracy cDNA sequence. These data provide both convincing evidence that the insertion of Tn5 into cDNA clones is sufficiently random for its effective use in large-scale cDNA sequencing as well as interesting insight about the sequence context preferred for insertion by Tn5.

Base Composition↗

Patome: a database server for biological sequence annotation and analysis in issued patents and published patent applications.

With the advent of automated and high-throughput techniques, the number of patent applications containing biological sequences has been increasing rapidly. However, they have attracted relatively little attention compared to other sequence resources. We have built a database server called Patome, which contains biological sequence data disclosed in patents and published applications, as well as their analysis information. The analysis is divided into two steps. The first is an annotation step in which the disclosed sequences were annotated with RefSeq database. The second is an association step where the sequences were linked to Entrez Gene, OMIM and GO databases, and their results were saved as a gene-patent table. From the analysis, we found that 55% of human genes were associated with patenting. The gene-patent table can be used to identify whether a particular gene or disease is related to patenting. Patome is available at http://www.patome.org/; the information is updated bimonthly.

Amino Acid Sequence↗

Gencube: centralized retrieval and integration of multi-omics resources from leading databases.

MOTIVATION: The volume of multi-omics data for diverse species is growing at an unprecedented rate, with new genome assemblies, related annotations, and high-throughput sequencing resources being submitted daily to various genomic data repositories. In response to this data influx, both existing and new databases are establishing optimized hierarchical structures to manage the vast amount of information. However, the lack of accessible command-line tools, combined with the functional limitations and unintuitive design of existing options, presents significant challenges for researchers. This gap underscores a critical need for a tool that enables streamlined retrieval and integration of omics data across these diverse repositories. RESULTS: We have developed Gencube, a command-line tool that enables centralized retrieval and integration of a comprehensive set of six different data types-genome assemblies, gene sets, annotations, sequences, comparative genomic data, and NGS-based omics resources-from various leading databases. AVAILABILITY AND IMPLEMENTATION: Gencube is a free and open-source tool, with its code available on GitHub: https://github.com/snu-cdrc/gencube and also archived on Zenodo: https://doi.org/10.5281/zenodo.14607649.

Databases, Genetic↗

Families of short interspersed elements in the genome of the oomycete plant pathogen, Phytophthora infestans.

The first known families of tRNA-related short interspersed elements (SINEs) in the oomycetes were identified by exploiting the genomic DNA sequence resources for the potato late blight pathogen, Phytophthora infestans. Fifteen families of tRNA-related SINEs, as well as predicted tRNAs, and other possible RNA polymerase III-transcribed sequences were identified. The size of individual elements ranges from 101 to 392 bp, representing sequences present from low (1) to highly abundant (over 2000) copy number in the P. infestans genome, based on quantitative PCR analysis. Putative short direct repeat sequences (6-14 bp) flanking the elements were also identified for eight of the SINEs. Predicted SINEs were named in a series prefixed infSINE (for infestans-SINE). Two SINEs were apparently present as multimers of tRNA-related units; four copies of a related unit for infSINEr, and two unrelated units for infSINEz. Two SINEs, infSINEh and infSINEi, were typically located within 400 bp of each other. These were also the only two elements identified as being actively transcribed in the mycelial stage of P. infestans by RT-PCR. It is possible that infSINEh and infSINEi represent active retrotransposons in P. infestans. Based on the quantitative PCR estimates of copy number for all of the elements identified, tRNA-related SINEs were estimated to comprise 0.3% of the 250 Mb P. infestans genome. InfSINE-related sequences were found to occur in species throughout the genus Phytophthora. However, seven elements were shown to be exclusive to P. infestans.

Base Sequence↗

Reliable identification of large numbers of candidate SNPs from public EST data.

High-resolution genetic analysis of the human genome promises to provide insight into common disease susceptibility. To perform such analysis will require a collection of high-throughput, high-density analysis reagents. We have developed a polymorphism detection system that uses public-domain sequence data. This detection system is called the single nucleotide polymorphism pipeline (SNPpipeline). The analytic core of the SNPpipeline is composed of three components: PHRED, PHRAP and DEMIGLACE. PHRED and PHRAP are components of a sequence analysis suite developed to perform the semi-automated analysis required for large-scale genomes (provided courtesy of P. Green). Using these informatics tools, which examine redundant raw expressed sequence tag (EST) data, we have identified more than 3,000 candidate single-nucleotide polymorphisms (SNPs). Empiric validation studies of a set of 192 candidates indicate that 82% identify variation in a sample of ten Centre d'Etudes Polymorphism Humain (CEPH) individuals. Our results suggest that existing sequence resources may serve as a valuable source for identifying genetic variation.

Algorithms↗

The EMBL Nucleotide Sequence Database.

The EMBL Nucleotide Sequence Database is a comprehensive database of DNA and RNA sequences directly submitted from researchers and genome sequencing groups and collected from the scientific literature and patent applications. In collaboration with DDBJ and GenBank the database is produced, maintained and distributed at the European Bioinformatics Institute (EBI) and constitutes Europe's primary nucleotide sequence resource. Database releases are produced quarterly and are distributed on CD-ROM. EBI's network services allow access to the most up-to-date data collection via Internet and World Wide Web interface, providing database searching and sequence similarity facilities plus access to a large number of additional databases.

Academies and Institutes↗

Sequence search algorithm assessment and testing toolkit (SAT).

MOTIVATION: The Sequence Search Algorithm Assessment and Testing Toolkit (SAT) aims to be a complete package for the comparison of different protein homology search algorithms. The structural classification of proteins can provide us with a clear criterion for judgment in homology detection. There have been several assessments based on structural sequences with classifications but a good deal of similar work is now being repeated with locally developed procedures and programs. The SAT will provide developers with a complete package which will save time and produce more comparable performance assessments for search algorithms. The package is complete in the sense that it provides a non-redundant large sequence resource database, a well-characterized query database of proteins domains, all the parsers and some previous results from PSI-BLAST and a hidden markov model algorithm. RESULTS: An analysis on two different data sets was carried out using the SAT package. It compared the performance of a full protein sequence database (RSDB100) with a non-redundant representative sequence database derived from it (RSDB50). The performance measurement indicated that the full database is sub-optimal for a homology search. This result justifies the use of much smaller and faster RSDB50 than RSDB100 for the SAT. AVAILABILITY: A web site is up. The whole packa ge is accessible via www and ftp. ftp://ftp.ebi.ac.uk/pub/contrib/jong/SAT http://cyrah.ebi.ac.uk:1111/Proj/Bio/SAT http://www.mrc-lmb.cam.ac.uk/genomes/SAT In the package, some previous assessment results produced by the package can also be found for reference. CONTACT: jong@ebi.ac.uk

Algorithms↗

Comparing gene expression profiles in human liver, gastric, and pancreatic tissues using full-length-enriched cDNA libraries.

In the post-genome-sequencing era, full-length cDNA-sequence resources are extremely useful for functional analyses of genes. In addition, comprehensive gene profiling of human tissues at the mRNA level is also useful in understanding the molecular mechanisms of tissue-specific functions and disease pathogenesis. In this study, to obtain a wide variety of full-length cDNA clones derived from digestive tissues, numerous expressed sequence tags were generated from libraries enriched with full-length cDNAs. In total, 13575 sequences were obtained from three cDNA libraries, which were constructed from tissues and cell lines of human liver, stomach, and pancreas. The integration of overlapping clones categorized the sequences into 5936 clusters (1666, 2746, and 2222 clusters in the liver, stomach, and pancreas, respectively). Of these, 1138 clones were scored as full-length cDNAs. Surprisingly, the redundant clones from all three tissues were assembled to show that only 101 genes (1.7% of the assembled 5936 genes) were shared. These results suggest that functional differences between tissues are probably related to their divergent gene expression profiles, and form a basis for understanding the molecular mechanisms underlying tissue-specific pathogenesis that are expressed in different organs. In addition, the full-length cDNAs obtained in this study should prove useful for future functional analyses of the genes expressed in digestive tissues.

Journal Article↗

The expanding role of microarrays in the investigation of macrophage responses to pathogens.

In the last few years, microarray technology has emerged as the method of choice for large-scale gene expression studies. It provides an efficient and rapid method to investigate the entire transcriptome of a cell. No research field has benefited more from microarray technology than the study of the exquisite interplay between pathogens and hosts. Numerous microarray studies have now been published in this field, which have provided insights into the mechanisms of host defence and the tactics employed by pathogens to circumvent these protection strategies. These studies have led to a more comprehensive understanding of the host immune response and identified new avenues of research for potential control strategies against pathogens. In the past, research has concentrated on human and mouse microarrays to investigate host-pathogen interactions, regardless of the host species. This trend is changing with the ever-expanding sequence resources now available for many pathogen and host species, including livestock animals. The use of species-specific microarrays has furthered our understanding of host-pathogen interactions for particular organisms and aided in the annotation of unknown genes. Macrophages play a central role in the host's innate and adaptive immune responses to pathogens. These cells are in the first line of defence and interact with a wide range of pathogens; many of which have evolved strategies to circumvent the macrophage defence mechanisms and survive within these cells. In this report, we review the wealth of studies using microarray technology to investigate the response of macrophages to pathogens. These studies illustrate how microarray technology has expanded our understanding of the dialogue between macrophage and pathogen and provide examples of the benefits and pitfalls of using this technique. Furthermore, we discuss the resources available to use microarray analysis to study the immune response of a non-human, non-rodent species, the cow.

Animals↗

Informatics-based learning resources for patients and their relatives in recovery.

In this paper we describe experiences from design of an informatics-based learning resource for patients and relatives. The prototype, REPARERE (learning REsource for PAtients and RElatives during REcovery), aims to support patients and their family recovering from heart surgery in meeting challenges in to daily living post discharge. Using recovery experiences and patient teaching material, REPARERE includes examples of textual information, video-clips, images and illustrations relevant to the recovery trajectory and a user's digitally represented profile. The development of the prototype focuses on flexibility and usability, tailoring and sequencing resources, and inclusion of recommendations for universal access. Development of web-based learning resources allows for exploration of 'just-in-case' and 'just-in-time' strategies to information retrieval and knowledge construction in health and learning trajectories. Findings from the literature, discussions with patients as well as health care providers indicate that unfulfilled information needs in the recovery period are common. Resources like REPARERE would be valuable supplement to facilitate patient learning about symptom management, self-care and coping while recovering.

Family↗

Antibodies against Epstein-Barr nuclear antigen (EBNA) in multiple sclerosis CSF, and two pentapeptide sequence identities between EBNA and myelin basic protein.

The Epstein-Barr virus (EBV) causes infectious mononucleosis and is linked to several disparate malignancies. Prior studies on patients with multiple sclerosis (MS) showed that 100% are EBV-seropositive and that their blood contains higher antibody titers than those of controls to both transformation and lytic cycle antigens. We performed three different assays for antibodies in CSF to three major EBV antigens from patients with MS and controls. Among 93 patients with MS, 79 (85%) had CSF that reacted with a 70 kD protein, shown to be the nuclear antigen, EBNA-1, whereas only 11 (13%) of 81 EBV-seropositive controls reacted, p less than 0.001. The CSF of all 14 MS patients, unreactive on immunoblots, contained oligoclonal bands on agarose electrophoresis. Together, the two techniques exhibit 100% sensitivity in the confirmatory diagnosis of MS. We also performed amino acid searches of the Protein Identification Resource sequence database for protein homologies to EBNA. Two pentapeptide identities were found between EBNA-1 and myelin basic protein: QKRPS and PRHRD. None of more than 32,000 other proteins in the database contained both pentapeptides. In healthy EBV-seropositive persons, the EBV-specific, MHC-restricted T lymphocytes keep the EBV-containing B lymphocytes locked in the transformed state. However, in the host genetically susceptible to MS, the same population of lymphocytes might recognize and interact with either of the two identified pentapeptides, inadvertently damaging MBP.

Amino Acid Sequence↗