Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequencing Resource”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

The genome sequence of silkworm, Bombyx mori.

We performed threefold shotgun sequencing of the silkworm (Bombyx mori) genome to obtain a draft sequence and establish a basic resource for comprehensive genome analysis. By using the newly developed RAMEN assembler, the sequence data derived from whole-genome shotgun (WGS) sequencing were assembled into 49,345 scaffolds that span a total length of 514 Mb including gaps and 387 Mb without gaps. Because the genome size of the silkworm is estimated to be 530 Mb, almost 97% of the genome has been organized in scaffolds, of which 75% has been sequenced. By carrying out a BLAST search for 50 characteristic Bombyx genes and 11,202 non-redundant expressed sequence tags (ESTs) in a Bombyx EST database against the WGS sequence data, we evaluated the validity of the sequence for elucidating the majority of silkworm genes. Analysis of the WGS data revealed that the silkworm genome contains many repetitive sequences with an average length of <500 bp. These repetitive sequences appear to have been derived from truncated transposons, which are interspersed at 2.5- to 3-kb intervals throughout the genome. This pattern suggests that silkworm may have an active mechanism that promotes removal of transposons from the genome. We also found evidence for insertions of mitochondrial DNA fragments at 9 sites. A search for Bombyx orthologs to Drosophila genes controlling sex determination in the WGS data revealed 11 Bombyx genes and suggested that the sex-determining systems differ profoundly between the two species.

Animals↗

A search tool for identification and analysis of conserved sequence patterns in Saccharomyces spp. orthologous promoter.

We describe a web-based resource to identify, search and analyze sequence patterns conserved in the multiple sequence alignments of orthologous promoters from closely related / distant Saccharomyces spp. The webtool interfaces with a database where conserved sequence patterns (greater than 4 bp) have been previously extracted from genome-wide promoter alignments, allowing one to carry out user-defined genome-wide searches for conserved sequences to assist in the discovery of novel promoter elements based on comparative genomics. The web-based server can be accessed at http://www2.imtech.res.in/ anand/sacch_prom_pat.html.

Base Sequence↗

Mapping PDB chains to UniProtKB entries.

MOTIVATION: UniProtKB/SwissProt is the main resource for detailed annotations of protein sequences. This database provides a jumping-off point to many other resources through the links it provides. Among others, these include other primary databases, secondary databases, the Gene Ontology and OMIM. While a large number of links are provided to Protein Data Bank (PDB) files, obtaining a regularly updated mapping between UniProtKB entries and PDB entries at the chain or residue level is not straightforward. In particular, there is no regularly updated resource which allows a UniProtKB/SwissProt entry to be identified for a given residue of a PDB file. RESULTS: We have created a completely automatically maintained database which maps PDB residues to residues in UniProtKB/SwissProt and UniProtKB/trEMBL entries. The protocol uses links from PDB to UniProtKB, from UniProtKB to PDB and a brute-force sequence scan to resolve PDB chains for which no annotated link is available. Finally the sequences from PDB and UniProtKB are aligned to obtain a residue-level mapping. AVAILABILITY: The resource may be queried interactively or downloaded from http://www.bioinf.org.uk/pdbsws/.

Amino Acid Sequence↗

Universal rule for coding sequence construction: TA/CG deficiency-TG/CT excess.

Each coding sequence is a finite resource as to the number and composition of four bases. Accordingly, the excessive recurrence of one base oligomer entails the noticeable underrepresentation by the other, so that if the former is the same in most, if not all, of the coding sequences, the latter too must necessarily be the same in all. Indeed, a previous series of studies on 20-odd divergent coding sequences established CTG as one of the most frequently recurring base trimers (if not the most frequent), and this excess was compensated by the underrepresentation by CG and TA dimer-containing base trimers. In this study, I have analyzed three additional coding sequences and reanalyzed one previously studied coding sequence. These four, derived from man, a plant, and a fish, were of variously lopsided base compositions that were not at all conducive to high recurrences of either CT dimer or CT and TG. Yet, the excess of CT and TG dimers accompanied by complementary deficiency of CG and TA dimers emerged as the common rule. Thus, I propose the above as the universal rule of coding sequence construction. The underrepresentation by CG and TA dimers within coding sequences explains why regulatory signals in intergenic spacers are of two kinds: one, TA dimer rich; and the other, CG dimer rich.

Adenine↗

Automated SNP detection in expressed sequence tags: statistical considerations and application to maritime pine sequences.

We developed an automated pipeline for the detection of single nucleotide polymorphisms (SNPs) in expressed sequence tag (EST) data sets, by combining three DNA sequence analysis programs: Phred, Phrap and PolyBayes. This application requires access to the individual electrophoregram traces. First, a reference set of 65 SNPs was obtained from the sequencing of 30 gametes in 13 maritime pine (Pinus pinaster Ait.) gene fragments (6671 bp), resulting in a frequency of 1 SNP every 102.6 bp. Second, parameters of the three programs were optimized in order to retrieve as many true SNPs, while keeping the rate of false positive as low as possible. Overall, the efficiency of detection of true SNPs was 83.1%. However, this rate varied largely as a function of the rare SNP allele frequency: down to 41% for rare SNP alleles (frequency < 10%), up to 98% for allele frequencies above 10%. Third, the detection method was applied to the 18498 assembled maritime pine (Pinus pinaster Ait.) ESTs, allowing to identify a total of 1400 candidate SNPs, in contigs containing between 4 and 20 sequence reads. These genetic resources, described for the first time in a forest tree species, were made available at http://www.pierroton.inra/genetics/Pinesnps. We also derived an analytical expression for the SNP detection probability as a function of the SNP allele frequency, the number of haploid genomes used to generate the EST sequence database, and the sample size of the contigs considered for SNP detection. The frequency of the SNP allele was shown to be the main factor influencing the probability of SNP detection.

Algorithms↗

Web services and workflow management for biological resources.

BACKGROUND: The completion of the Human Genome Project has resulted in large quantities of biological data which are proving difficult to manage and integrate effectively. There is a need for a system that is able to automate accesses to remote sites and to "understand" the information that it is managing in order to link data properly. Workflow management systems combined with Web Services are promising Information and Communication Technologies (ICT) tools. Some have already been proposed and are being increasingly applied to the biomedical domain, especially as many biology-related Web Services are now becoming available. Information on biological resources and on genomic sequences mutations are two examples of very specialized datasets that are useful for specific research domains. RESULTS: The architecture of a system that is able to access and execute predefined workflows is presented in this paper. Web Services allowing access to the IARC TP53 Mutation Database and CABRI catalogues of biological resources have been implemented and are available on-line. Example workflows which retrieve data from these Web Services have also been created and are available on-line. CONCLUSION: We present a general architecture and some building blocks for the implementation of a system that is able to remotely execute workflows of biomedical interest and show how this approach can effectively produce useful outputs. The further development and implementation of Web Services allowing access to an exhaustive set of biomedical databases and the creation of effective and useful workflows will improve the automation of in-silico analysis.

Animals↗

Microbial genomes and "missing" enzymes: redefining biochemical pathways.

A biochemical pathway is the representation of a defined set of substrates, enzyme reactions and products linked together to generate an outcome beneficial to a living cell. Microbial genome sequence data are unparalleled resources for understanding cellular metabolism without the prior definitions imposed by classical biochemistry. Simple analysis of three well-studied biochemical pathways (the tricarboxylic acid cycle, pentose phosphate pathway and glycolysis) from the 17 publicly available microbial genomes has shown that these pathways may rarely occur as previously defined. Therefore, following whole-genome sequencing it has become necessary to redefine the "classical" biochemical steps leading from substrate to end-product for each pathway. Often, unique or alternative reactions appear to be required in order to maintain pathway functionality where expected enzyme reactions (as defined by the presence or absence of the corresponding genes) are "missing". Conversely, such enzymes may be accounted for by: (1) the presence of low sequence similarity or novel genes encoding enzymes performing the same or similar functions, (2) the presence of multienzyme proteins, (3) incorrectly assigned gene identities in genome databases, and (4) known enzyme functions that have yet to be correlated with a gene sequence. Most importantly, the presence of a gene sequence does not necessarily ensure that its corresponding enzyme is actually functional. This may be due to the presence of nonactive remnant genes, evolutionary pressures leading to loss of function, inactivating mutations, the lack of transcription/translation, and post-translational processing. Modifications at the gene and/or functional levels, as well as the possible use of alternative enzymes, must be considered when reconstructing biochemical pathways for fully sequenced microbial genomes.

Bacteria↗

A statewide instructional television program via satellite for RN-to-BSN students.

Instructional television (ITV) is providing the bridge between nursing educational needs and health care resources, allowing nurses in rural areas to attend a university without leaving their jobs or relocating. This article describes the experiences of developing a baccalaureate degree completion program for nursing via satellite. To create a successful distance education program, it was necessary to re-examine curricular structure, objectives, course descriptions, the basis of course sequences, teaching strategies, library resources, and learner needs. It was vital to modify the student and faculty support services on campus for use with distance learners.

Education, Nursing, Baccalaureate↗

COmplete GENome Tracking (COGENT): a flexible data environment for computational genomics.

SUMMARY: We present a database of fully sequenced and published genomes to facilitate the re-distribution of data and ensure reproducibility of results in the field of computational genomics. For its design we have implemented an extremely simple yet powerful schema to allow linking of genome sequence data to other resources. AVAILABILITY: http://maine.ebi.ac.uk:8000/services/cogent/

Computational Biology↗

A 2-megabase physical contig incorporating 43 DNA markers on the human X chromosome at p11.23-p11.22 from ZNF21 to DXS255.

A comprehensive physical contig of yeast artificial chromosomes (YACs) and cosmid clones between ZNF21 and DXS255 has been constructed, spanning 2 Mb within the region Xp11.23-p11.22. As a portion of the region was found to be particularly unstable in yeast, the integrity of the contig is dependent on additional information provided by the sequence-tagged site (STS) content of cosmid clones and DNA marker retention in conventional and radiation hybrids. The contig was formatted with 43 DNA markers, including 19 new STSs from YAC insert ends and an internal Alu-PCR product. The density of STSs across the contig ranges from one marker every 20 kb to one every 60 kb, with an average density of one marker every 50 kb. The relative order of previously known genes and expressed sequence tags in this region is predicted to be Xpter-ZNF21-DXS7465E (MG66)-DXS7927E (MG81)-WASP, DXS1011E, DXS7467E (MG21)-DXS- 7466E (MG44)-GATA1-DXS7469E (Xp664)-TFE3-SYP (DXS1007E)-Xcen. This contig extends the coverage in Xp11 and provides a framework for the future identification and mapping of new genes, as well as the resources for developing DNA sequencing templates.

Animals↗

In silico sequence evolution with site-specific interactions along phylogenetic trees.

MOTIVATION: A biological sequence usually has many sites whose evolution depends on other positions of the sequence, but this is not accounted for by commonly used models of sequence evolution. Here we introduce a Markov model of nucleotide sequence evolution in which the instantaneous substitution rate at a site depends on the states of other sites. Based on the concept of neighbourhood systems, our model represents a universal description of arbitrarily complex dependencies among sites. RESULTS: We show how to define complex models for some illustrative examples and demonstrate that our method provides a versatile resource for simulations of sequence evolution with site-specific interactions along a tree. For example, we are able to simulate the evolution of RNA taking into account both secondary structure as well as pseudoknots and other tertiary interactions. To this end, we have developed a program Simulating Site-Specific Interactions (SISSI) that simulates evolution of a nucleotide sequence along a phylogenetic tree incorporating user defined site-specific interactions. Furthermore, our method allows to simulate more complex interactions among nucleotide and other character based sequences. AVAILABILITY: We implemented our method in an ANSI C program SISSI which runs on UNIX/Linux, Windows and Mac OS systems, including Mac OS X. SISSI is available at http://www.bi.uni-duesseldorf.de/software/sissi/

Algorithms↗

Construction and characterization of partial digest DNA libraries made from flow-sorted human chromosome 16.

In this report, we present the techniques used for the construction of chromosome-specific partial digest libraries from flow-sorted chromosomes and the characterization of two such libraries from human chromosome 16. These libraries were constructed to provide materials for use in the development of a high-resolution physical map of human chromosome 16, and as part of a distributive effort on the National Laboratory Gene Library Project. Libraries with 20-fold coverage were made in Charon-40 (LA16NL03) and in sCos-1 (LA16NC02) after chromosome 16 was sorted from a mouse-human monochromosomal hybrid cell line containing a single homologue of human chromosome 16. Both libraries are approximately 90% enriched for human chromosome 16, have low nonrecombinant backgrounds, and are highly representative for human chromosome-16 sequences. The cosmid library in particular has provided a valuable resource for the isolation of coding sequences, and in the ongoing development of a physical map of human chromosome 16.

Animals↗

Mitogenomic exploration of higher teleostean phylogenies: a case study for moderate-scale evolutionary genomics with 38 newly determined complete mitochondrial DNA sequences.

Although adequate resolution of higher-level relationships of organisms apparently requires longer DNA sequences than those currently being analyzed, limitations of time and resources present difficulties in obtaining such sequences from many taxa. For fishes, these difficulties have been overcome by the development of a PCR-based approach for sequencing the complete mitochondrial genome (mitogenome), which employs a long PCR technique and many fish-versatile PCR primers. In addition, recent studies have demonstrated that such mitogenomic data are useful and decisive in resolving persistent controversies over higher-level relationships of teleosts. As a first step toward resolution of higher teleostean relationships, which have been described as the "(unresolved) bush at the top of the tree," we investigated relationships using mitogenomic data from 48 purposefully chosen teleosts, of which those from 38 were newly determined during the present study (a total of 632,315 bp), using the above method. Maximum-parsimony and maximum-likelihood analyses were conducted with the data set that comprised concatenated nucleotide sequences from 12 protein-coding genes (excluding the ND6 gene and third codon positions) and 22 transfer RNA (tRNA) genes (stem regions only) from the 48 species. The resultant two trees from the two methods were well resolved and largely congruent, with many internal branches supported by high statistical values. The tree topologies themselves, however, exhibited considerable variation from the previous morphology-based cladistic hypotheses, with most of the latter being confidently rejected by the mitogenomic data. Such incongruence resulted largely from the phylogenetic positions or limits of long-standing problematic taxa, which were quite unexpected from previous morphological and molecular analyses. We concluded that the present study provided a basis of and guidelines for future investigations of teleostean evolutionary mitogenomics and that purposeful higher-density taxonomic sampling, subsequent sequencing efforts, and phylogenetic analyses of their mitogenomes may be decisive in resolving persistent controversies over higher-level relationships of teleosts, the most diversified group of all vertebrates, comprising over 23,500 extant species.

Animals↗

The IMAGE project: methodological issues for the molecular genetic analysis of ADHD.

The genetic mechanisms involved in attention deficit hyperactivity disorder (ADHD) are being studied with considerable success by several centres worldwide. These studies confirm prior hypotheses about the role of genetic variation within genes involved in the regulation of dopamine, norepinephrine and serotonin neurotransmission in susceptibility to ADHD. Despite the importance of these findings, uncertainties remain due to the very small effects sizes that are observed. We discuss possible reasons for why the true strength of the associations may have been underestimated in research to date, considering the effects of linkage disequilibrium, allelic heterogeneity, population differences and gene by environment interactions. With the identification of genes associated with ADHD, the goal of ADHD genetics is now shifting from gene discovery towards gene functionality--the study of intermediate phenotypes ('endophenotypes'). We discuss methodological issues relating to quantitative genetic data from twin and family studies on candidate endophenotypes and how such data can inform attempts to link molecular genetic data to cognitive, affective and motivational processes in ADHD. The International Multi-centre ADHD Gene (IMAGE) project exemplifies current collaborative research efforts on the genetics of ADHD. This European multi-site project is well placed to take advantage of the resources that are emerging following the sequencing of the human genome and the development of international resources for whole genome association analysis. As a result of IMAGE and other molecular genetic investigations of ADHD, we envisage a rapid increase in the number of identified genetic variants and the promise of identifying novel gene systems that we are not currently investigating, opening further doors in the study of gene functionality.

Journal Article↗

Development of a panel of monochromosomal somatic cell hybrids for rapid gene mapping.

We have assembled a panel of monochromosomal somatic cell hybrids for use in gene mapping. DNA from each individual hybrid was used as a probe on normal human metaphases to identify the human chromosome and any fragments by reverse painting. To test the efficiency of the panel PCR amplification of DNA from the monochromosomal somatic cell hybrid panel was used in combination with human specific oligonucleotide primers to assign alpha-catenin (CTNNA1) and p21/WAF1 to chromosomes 5 and 6 respectively. These genes were localized further using hybrids containing specific translocations to 5q11-qter and 6p21 respectively. We also developed primers to enable us to assign 17 ESTs sequenced by the HGMP Resource Centre. The hybrid panel was developed with support of the UK HGMP and the DNA is available to all registered users.

Base Sequence↗

Dopamine signaling in Caenorhabditis elegans-potential for parkinsonism research.

The nematode Caenorhabditis elegans is an attractive model system for the study of many biological processes. It possesses a simple nervous system with known anatomy and connectivity, is conveniently and cheaply cultured in the laboratory, and is amenable to many genetic manipulations that are impossible in mammalian systems. The recent completion of the C. elegans genome sequence provides a rich resource of genomic and bioinformatic data to researchers in diverse fields. This organism, however, has been underexploited in the studies of many basic processes related to nervous system function, neuropsychiatric disorders and neuromuscular function. Anatomical, biochemical, behavioral, pharmacological and genetic evidence accumulated to date strongly suggests that dopamine is used as a neurotransmitter by C. elegans, and that its effects are mediated through pathway(s) that share many features with those of mammals. DNA sequence analysis reveals genes highly homologous to those encoding mammalian dopamine receptors. Probably, C. elegans has dopamine receptors that transduce environmental cues into behaviors, and these receptors pharmacologically most closely resemble the D2 family. Here we present a review of the current state of research into the dopamine system of the worm, focussing on its potential for use in the study of biological processes related to parkinsonism.

Journal Article↗

Digital cloning: identification of human cDNAs homologous to novel kinases through expressed sequence tag database searching.

Identification of novel kinases based on their sequence conservation within kinase catalytic domain has relied so far on two major approaches, low-stringency hybridization of cDNA libraries, and PCR method using degenerate primers. Both of these approaches at times are technically difficult and time-consuming. We have developed a procedure that can significantly reduce the time and effort involved in searching for novel kinases and increase the sensitivity of the analysis. This procedure exploits the computer analysis of a vast resource of human cDNA sequences represented in the expressed sequence tag (EST) database. Seventeen novel human cDNA clones showing significant homology to serine/threonine kinases, including STE-20, CDK- and YAK-related family kinases, were identified by searching EST database. Further sequence analysis of these novel kinases obtained either directly from EST clones or from PCR-RACE products confirmed their identity as protein kinases. Given the rapid accumulation of the EST database and the advent of powerful computer analysis software, this approach provides a fast, sensitive, and economical way to identify novel kinases as well as other genes from EST database.

Amino Acid Sequence↗