Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequencing Resource”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

The Schistosoma mansoni gene index: gene discovery and biology by reconstruction and analysis of expressed gene sequences.

Expressed sequence tag (EST) sequencing and analysis is a primary research tool to identify and characterize the Schistosoma mansoni transcriptome. As part of our gene discovery effort, a total of 5,793 ESTs have been generated from clones selected randomly from complementary DNA (cDNA) libraries constructed from male and female adult worms. Assembly analysis of all the 16,813 public S. mansoni ESTs has identified 1,920 distinct tentative consensus sequences (TCs) and 5,571 nonoverlapping ESTs (singletons). Of these, 376 TCs (20%) and 1,449 singletons (26%) are unique to the SUNY/TIGR sequencing effort. Tentative consensus sequences and singletons were distributed into various categories of biological roles associated with cell structure, metabolism, protein fate, signal transduction, transcription, protein synthesis, transporters, and cell growth. The TCs and singletons represent transcripts that can be used as a resource for functional annotation of genomic sequence data, comparative sequence analysis, and cDNA clone selection for microarray projects. The utility of EST analysis is demonstrated by identifying new protease genes, which may be involved in hemoglobin degradation.

Amino Acid Sequence↗

The Bradyrhizobium japonicum serocluster 123 hyperreiterated DNA region, HRS1, has DNA and amino acid sequence homology to IS1380, an insertion sequence from Acetobacter pasteurianus.

We have sequenced and analyzed the hyperreiterated DNA region, HRS1, from Bradyrhizobium japonicum USDA 424. The 2.1-kb HRS1 fragment is closely linked to the B. japonicum common and genotype-specific nodulation genes in serogroup 123 and 127 strains. Southern hybridization analyses indicated that one copy of HRS1 is also located next to the fixRnifA locus in B. japonicum USDA 424. Nucleotide sequence analysis revealed the presence of a 4-bp target site duplication in HRS1 which is identical to a terminal repeat found in the B. japonicum USDA 110 repeated sequence RS alpha. Computer searches of the PIR (Protein Identification Resource) protein data base revealed a high degree of amino acid sequence homology between a putative 329-amino-acid polypeptide from HRS1 and a large polypeptide from IS1380, an insertion sequence from Acetobacter pasteurianus. RNA slot blot hybridizations suggest that transcripts showing homology to HRS1 are constitutively produced in strains USDA 424 (serogroup 127) and USDA 438 (serogroup 123).

Acetobacter↗

A first generation physical map of the medaka genome in BACs essential for positional cloning and clone-by-clone based genomic sequencing.

In order to realize the full potential of the medaka as a model system for developmental biology and genetics, characterized genomic resources need to be established, culminating in the sequence of the medaka genome. To facilitate the map-based cloning of genes underlying induced mutations and to provide templates for clone-based genomic sequencing, we have created a first-generation physical map of the medaka genome in bacterial artificial chromosome (BAC) clones. In particular, we exploited the synteny to the closely related genome of the pufferfish, Takifugu rubripes, by marker content mapping. As a first step, we clustered 103,144 public medaka EST sequences to obtain a set of 21,121 non-redundant sequence entities. Avoiding oversampling of gene-dense regions, 11,254 of EST clusters were successfully matched against the draft sequence of the fugu genome, and 2363 genes were selected for the BAC map project. We designed 35mer oligonucleotide probes from the selected genes and hybridized them against 64,500 BAC clones of strains Cab and Hd-rR, representing 14-fold coverage of the medaka genome. Our data set is further supplemented with 437 results generated from PCR-amplified inserts of medaka cDNA clones and BAC end-fragment markers. Our current, edited, first generation medaka BAC map consists of 902 map segments that cover about 74% of the medaka genome. The map contains 2721 markers. Of these, 2534 are from expressed sequences, equivalent to a non-redundant set of 2328 loci. The 934 markers (724 different) are anchored to the medaka genetic map. Thus, genetic map assignments provide immediate access to underlying clones and contigs, simplifying molecular access to candidate gene regions and their characterization.

Animals↗

The CATH database: an extended protein family resource for structural and functional genomics.

The CATH database of protein domain structures (http://www.biochem.ucl.ac.uk/bsm/cath_new) currently contains 34 287 domain structures classified into 1383 superfamilies and 3285 sequence families. Each structural family is expanded with domain sequence relatives recruited from GenBank using a variety of efficient sequence search protocols and reliable thresholds. This extended resource, known as the CATH-protein family database (CATH-PFDB) contains a total of 310 000 domain sequences classified into 26 812 sequence families. New sequence search protocols have been designed, based on these intermediate sequence libraries, to allow more regular updating of the classification. Further developments include the adaptation of a recently developed method for rapid structure comparison, based on secondary structure matching, for domain boundary assignment. The philosophy behind CATHEDRAL is the recognition of recurrent folds already classified in CATH. Benchmarking of CATHEDRAL, using manually validated domain assignments, demonstrated that 43% of domains boundaries could be completely automatically assigned. This is an improvement on a previous consensus approach for which only 10-20% of domains could be reliably processed in a completely automated fashion. Since domain boundary assignment is a significant bottleneck in the classification of new structures, CATHEDRAL will also help to increase the frequency of CATH updates.

Animals↗

MOsDB: an integrated information resource for rice genomics.

The MIPS Rice (Oryza sativa) database (MOsDB; http://mips.gsf.de/proj/rice) provides a comprehensive data collection dedicated to the genome information of rice. Rice (O. sativa L.) is one of the most important food crops for over half the world's population and serves as a major model system in cereal genome research. MOsDB integrates data from two publicly available rice genomic sequences, O. sativa L. ssp. indica and O. sativa L. ssp. japonica. Besides regularly updated rice genome sequence information, MOsDB provides an integrated resource for associated analysis data, e.g. internal and external annotation information as well as a complex characterization of all annotated rice genes. The MOsDB web interface supports various search options and allows browsing the database content. MOsDB is continuously expanding to include an increasing range of data type and the growing amount of information on the rice genome.

DNA Transposable Elements↗

Uprobe: a genome-wide universal probe resource for comparative physical mapping in vertebrates.

Interspecies comparisons are important for deciphering the functional content and evolution of genomes. The expansive array of >70 public vertebrate genomic bacterial artificial chromosome (BAC) libraries can provide a means of comparative mapping, sequencing, and functional analysis of targeted chromosomal segments that is independent and complementary to whole-genome sequencing. However, at the present time, no complementary resource exists for the efficient targeted physical mapping of the majority of these BAC libraries. Universal overgo-hybridization probes, designed from regions of sequenced genomes that are highly conserved between species, have been demonstrated to be an effective resource for the isolation of orthologous regions from multiple BAC libraries in parallel. Here we report the application of the universal probe design principal across entire genomes, and the subsequent creation of a complementary probe resource, Uprobe, for screening vertebrate BAC libraries. Uprobe currently consists of whole-genome sets of universal overgo-hybridization probes designed for screening mammalian or avian/reptilian libraries. Retrospective analysis, experimental validation of the probe design process on a panel of representative BAC libraries, and estimates of probe coverage across the genome indicate that the majority of all eutherian and avian/reptilian genes or regions of interest can be isolated using Uprobe. Future implementation of the universal probe design strategy will be used to create an expanded number of whole-genome probe sets that will encompass all vertebrate genomes.

Alligators and Crocodiles↗

OGRe: a relational database for comparative analysis of mitochondrial genomes.

Organellar Genome Retrieval (OGRe) is a relational database of complete mitochondrial genome sequences for over 250 Metazoan species. OGRe provides a resource for the comparative analysis of mitochondrial genomes at several levels. At the sequence level, OGRe allows the retrieval of any selected set of mitochondrial genes from any selected set of species. Species are classified using a taxonomic system that allows easy selection of related groups of species. Sequence alignments are also available for some species. At the level of individual nucleotides, the system contains information on base frequencies and codon usage frequencies that can be compared between organisms. At the level of whole genomes, OGRe provides several ways of visualizing information on gene order. Diagrams illustrating the genome arrangement can be generated for any selected set of species automatically from the information in the database. Searches can be done based on gene arrangement to find sets of species that have the same order as one another. Diagrams for pairwise comparison of species can be produced that show the positions of break-points in the gene order and use colour to highlight the sections of the genome that have moved. OGRe is available from http://www.bioinf.man.ac.uk/ogre.

Animals↗

Identifying conservation units within captive chimpanzee populations.

One of the primary objectives in the captive management of any endangered primate is to preserve as much as possible the genetic diversity that has evolved and still exists in wild gene pools. The rationale for this is based on the theoretical understanding of the relationship between genetic diversity and fitness in response to selection. There remains little consensus, however, as to the type of genetic data that should be used to monitor captive populations. In order to develop a deeper understanding of the degree and nature of genetic diversity among "wild" chimpanzee gene pools, as well as to determine if one type of genetic data is more useful than others, DNA sequence data were generated at three unlinked, nonrepetitive nuclear loci, one polymorphic microsatellite, and the mitochondrial D-loop for 59 unrelated common and pygmy chimpanzees. The results suggest that: 1) data from nuclear loci can be used to differentiate common chimpanzee subspecies; 2) pygmy chimpanzees may have less genetic diversity than common chimpanzees; 3) shared microsatellite alleles do not always indicate identity by descent; and 4) nonrepetitive loci provide unique insights into evolutionary relationships and provide useful information for captive management programs.

Animals↗

Phylogeography and conservation genetics of Eld's deer (Cervus eldi).

Eld's deer (Cervus eldi) is a highly endangered cervid, distributed historically throughout much of South Asia and Indochina. We analysed variation in the mitochondrial DNA (mtDNA) control region for representatives of all three Eld's deer subspecies to gain a better understanding of the genetic population structure and evolutionary history of this species. A phylogeny of mtDNA haplotypes indicates that the critically endangered and ecologically divergent C. eldi eldi is related more closely to C. e. thamin than to C. e. siamensis, a result that is consistent with biogeographic considerations. The results also suggest a strong degree of phylogeographic structure both between subspecies and among populations within subspecies, suggesting that dispersal of individuals between populations has been very limited historically. Haplotype diversity was relatively high for two of the three subspecies (thamin and siamensis), indicating that recent population declines have not yet substantially eroded genetic diversity. In contrast, we found no haplotype variation within C. eldi eldi or the Hainan Island population of C. eldi siamensis, two populations which are known to have suffered severe population bottlenecks. We also compared levels of haplotype and nucleotide diversity in an unmanaged captive population, a managed captive population and a relatively healthy wild population. Diversity indices were higher in the latter two, suggesting the efficacy of well-designed breeding programmes for maintaining genetic diversity in captivity. Based on significant genetic differentiation among Eld's deer subspecies, we recommend the continued management of this species in three distinct evolutionarily significant units (ESUs). Where possible, it may be advisable to translocate individuals between isolated populations within a subspecies to maintain levels of genetic variation in remaining Eld's deer populations.

Animal Population Groups↗

SIMAP: the similarity matrix of proteins.

Similarity Matrix of Proteins (SIMAP) (http://mips.gsf.de/simap) provides a database based on a pre-computed similarity matrix covering the similarity space formed by >4 million amino acid sequences from public databases and completely sequenced genomes. The database is capable of handling very large datasets and is updated incrementally. For sequence similarity searches and pairwise alignments, we implemented a grid-enabled software system, which is based on FASTA heuristics and the Smith-Waterman algorithm. Our ProtInfo system allows querying by protein sequences covered by the SIMAP dataset as well as by fragments of these sequences, highly similar sequences and title words. Each sequence in the database is supplemented with pre-calculated features generated by detailed sequence analyses. By providing WWW interfaces as well as web-services, we offer the SIMAP resource as an efficient and comprehensive tool for sequence similarity searches.

Databases, Protein↗

Mutation detection using mass spectrometric separation of tiny oligonucleotide fragments.

A DNA mutation detection protocol able to identify and characterize a previously unknown change in a given sequence in a rapid, efficient, sensitive, and inexpensive manner is required to take advantage of the resources now available to researchers through the genome sequencing projects. We have developed a method based on base-specific cleavage of polymerase chain reaction (PCR) products and then separation of the fragments by matrix-assisted laser desorption ionization-mass spectrometry (MALDI-MS), which can meet these criteria. Differences are seen as the presence, absence, or mass change of peaks corresponding to fragments affected by the base difference. This technique is shown through the detection of a polymorphism in the 3' untranslated region of IL12p40 from a double-stranded PCR product, and the detection of a single nucleotide polymorphism between two mouse strains. The sensitivity of the technique can be increased with the use of postsource decay, which enables differentiation of two fragments of identical mass but different sequence. The level of specificity and the rapid sample analysis time lend this technique to the mass screening of individuals for sequence changes and, in combination with MS sequencing methods, could be used to facilitate rapid resequencing of DNA.

3' Untranslated Regions↗

Genome annotation: from sequence to biology.

The genome sequence of an organism is an information resource unlike any that biologists have previously had access to. But the value of the genome is only as good as its annotation. It is the annotation that bridges the gap from the sequence to the biology of the organism. The aim of high-quality annotation is to identify the key features of the genome - in particular, the genes and their products. The tools and resources for annotation are developing rapidly, and the scientific community is becoming increasingly reliant on this information for all aspects of biological research.

Base Sequence↗

Using the chicken genome sequence in the development and mapping of genetic markers in the turkey (Meleagris gallopavo).

The efficacy of employing the chicken genome sequence in developing genetic markers and in mapping the turkey genome was studied. Eighty previously uncharacterized microsatellite markers were identified for the turkey using BLAST alignment to the chicken genome. The chicken sequence was then used to develop primers for polymerase chain reaction where the turkey sequence was either unavailable or insufficient. A total of 78 primer sets were tested for amplification and polymorphism in the turkey, and informative markers were genetically mapped. Sixty-five (83%) amplified turkey genomic DNA, and 33 (42%) were polymorphic in the University of Minnesota/Nicholas Turkey Breeding Farms mapping families. All but one marker genetically mapped to the position predicted from the chicken genome sequence. These results demonstrate the usefulness of the chicken sequence for the development of genomic resources in other avian species.

Alleles↗

Development and testing of a high-density cDNA microarray resource for cattle.

A cDNA microarray resource has been developed with the goal of providing integrated functional genomics resources for cattle. The National Bovine Functional Genomics Consortium's (NBFGC) expressed sequence tag (EST) collection was established in 2001 to develop resources for functional genomics research. The NBFGC EST collection and microarray contains 18,263 unique transcripts, derived from many different tissue types and various physiologically important states within these tissues. The NBFGC microarray has been tested for false-positive rates using self-self hybridizations and was shown to yield robust results in test microarray experiments. A web-accessible database has been established to provide pertinent data related to NBFGC clones, including sequence data, BLAST results, and ontology information. The NBFGC microarray represents the largest cDNA microarray for a livestock species prepared to date and should prove to be a valuable tool in studying genome-wide gene expression in cattle.

Animals↗

A high level interface to SCOP and ASTRAL implemented in python.

BACKGROUND: Benchmarking algorithms in structural bioinformatics often involves the construction of datasets of proteins with given sequence and structural properties. The SCOP database is a manually curated structural classification which groups together proteins on the basis of structural similarity. The ASTRAL compendium provides non redundant subsets of SCOP domains on the basis of sequence similarity such that no two domains in a given subset share more than a defined degree of sequence similarity. Taken together these two resources provide a 'ground truth' for assessing structural bioinformatics algorithms. We present a small and easy to use API written in python to enable construction of datasets from these resources. RESULTS: We have designed a set of python modules to provide an abstraction of the SCOP and ASTRAL databases. The modules are designed to work as part of the Biopython distribution. Python users can now manipulate and use the SCOP hierarchy from within python programs, and use ASTRAL to return sequences of domains in SCOP, as well as clustered representations of SCOP from ASTRAL. CONCLUSION: The modules make the analysis and generation of datasets for use in structural genomics easier and more principled.

Database Management Systems↗

Complete set of ORF clones of Escherichia coli ASKA library (a complete set of E. coli K-12 ORF archive): unique resources for biological research.

Based on the genomic sequence data of Escherichia coli K-12 strain, we have constructed a complete set of cloned individual genes encoding Histidine-tagged proteins with or without GFP fused for functional genomic analysis. Each clone encodes a protein of predicted ORF attached by Histidines and seven spacer amino acids at the N-terminal end, and five spacer amino acids and GFP at the C-terminal end. SfiI restriction sites are generated at both the N- and C-terminal boundaries of ORF upon cloning, which enables easy transfer of ORF to other vector systems by cutting with SfiI. Expression of cloned ORF is under the control of an IPTG-inducible promoter, which is strictly repressed by lacI(q) repressor gene product. The set of cloned ORFs described here should provide unique resources for systematic functional genomic approaches including (i) construction of DNA microarray, (ii) production and purification of proteins, (iii) analysis of protein localization by monitoring GFP fluorescence and (iv) analysis of protein-protein interaction.

Escherichia coli K12↗

Sequence of the HindIII T fragment of human cytomegalovirus, which encodes a DNA helicase.

The DNA sequence of the HindIII T fragment of human cytomegalovirus strain AD169 has been determined. This 6225 bp sequence has been analysed for transcription signals and probable open reading frames. Similarities with herpes simplex virus, varicella-zoster virus and Epstein-Barr virus genes were observed for three of the predicted open reading frames; a virion protein and a unique DNA helicase are believed to be the functional products of two of these open reading frames. Two other open reading frames are novel in that no homologues could be found, either in the known herpesvirus sequences or in the Protein Identification Resource database. Both of these open reading frames also lie in the genomic coding region of a 5.0 kb RNA which is transcribed throughout the infectious cycle.

Amino Acid Sequence↗

Analyses of beta-1 syntrophin, syndecan 2 and gem GTPase as candidates for chicken muscular dystrophy.

Despite intensive studies of muscular dystrophy of chicken, the responsible gene has not yet been identified. Our recent studies mapped the genetic locus for abnormal muscle (AM) of chicken with muscular dystrophy to chromosome 2q using the Kobe University (KU) resource family, and revealed the chromosome region where the AM gene is located has conserved synteny to human chromosome 8q11-24.3, where the beta-1 syntrophin (SNTB1), syndecan 2 (SDC2) and Gem GTPase (GEM) genes are located. It is reasonable to assume those genes might be candidates for the AM gene. In this study, we cloned and sequenced the chicken SNTB1, SDC2 and GEM genes, and identified sequence polymorphisms between parents of the resource family. The polymorphisms were genotyped to place these genes on the chicken linkage map. The AM gene of chromosome 2q was mapped 130 cM from the distal end, and closely linked to calbindin 1 (CALB1). SNTB1 and SDC2 genes were mapped 88.5 cM distal and 27.6 cM distal from the AM gene, while the GEM gene was mapped 18.5 cM distal from the AM gene and 9.1 cM proximal from SDC2. Orthologues of SNTB1, SDC2 and GEM were syntenic to human chromosome 8q. SNTB1, SDC2 and GEM did not correspond to the AM gene locus, suggesting it is unlikely they are related to chicken muscular dystrophy. However, this result also suggests that the genes located in the proximal region of the CALB1 gene on human chromosome 8q are possible candidates for this disease.

Amino Acid Sequence↗