Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequencing Resource”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Mouse inbred strain sequence information and yin-yang crosses for quantitative trait locus fine mapping.

The shared ancestry of mouse inbred strains, together with the availability of sequence and phenotype information, is a resource that can be used to map quantitative trait loci (QTL). The difficulty in using only sequence information lies in the fact that in most instances the allelic state of the QTL cannot be unambiguously determined in a given strain. To overcome this difficulty, the performance of multiple crosses between various inbred strains has been proposed. Here we suggest and evaluate a general approach, which consists of crossing the two strains used initially to map the QTL and any new strain. We have termed these crosses "yin-yang," because they are complementary in nature as shown by the fact that the QTL will necessarily segregate in only one of the crosses. We used the publicly available SNP database of chromosome 16 to evaluate the mapping resolution achievable through this approach. Although on average the improvement of mapping resolution using only four inbred strains was relatively small (i.e., reduction of the QTL-containing interval by half at most), we found a great degree of variability among different regions of chromosome 16 with regard to mapping resolution. This suggests that with a large number of strains in hand, selecting a small number of strains may provide a significant contribution to the fine mapping of QTL.

Alleles↗

Association of INOS, TRAIL, TGF-beta2, TGF-beta3, and IgL genes with response to Salmonella enteritidis in poultry.

Several candidate genes were selected, based on their critical roles in the host's response to intracellular bacteria, to study the genetic control of the chicken response to Salmonella enteritidis (SE). The candidate genes were: inducible nitric oxide synthase (INOS), tumor necrosis factor related apoptosis inducing ligand (TRAIL), transforming growth factor beta2 (TGF-beta2), transforming growth factor beta3 (TGF-beta3), and immunoglobulin G light chain (IgL). Responses to pathogenic SE colonization or to SE vaccination were measured in the Iowa Salmonella response resource population (ISRRP). Outbred broiler sires and three diverse, highly inbred dam lines produced 508 F1 progeny, which were evaluated as young chicks for either bacterial load isolated from spleen or cecum contents after pathogenic SE inoculation, or the circulating antibody level after SE vaccination. Fragments of each gene were sequenced from the founder lines of the resource population to identify genomic sequence variation. Single nucleotide polymorphisms (SNP) were identified, then PCR-RFLP techniques were developed to genotype the F1 resource population. Linear mixed models were used for statistical analyses. Because the inbred dam lines always contributed one copy of the same allele, the heterozygous sire allele effects could be assessed in the F1 generation. Association analyses revealed significant effects of the sire allele of TRAIL-StyI on the spleen (P <0.07) and cecum (P <0.0002) SE bacterial load. Significant effects (P <0.04) were found on the cecum bacterial load for TGF-beta3-BsrI. Varied and moderate association was found for SE vaccine antibody response for all genes. This is the first reported study on the association of SNP in INOS, TRAIL, TGF-beta2, TGF-beta3, and IgL with the chicken response to SE. Identification of candidate genes to improve the immune response may be useful for marker-assisted selection to enhance disease resistance.

Animals↗

Drosophila-related expressed sequences.

The study of model organisms has been instrumental towards the elucidation of the basic mechanisms of human biology. Drosophila melanogaster has been the target of extensive genetic analyses over the past 90 years and a notable amount of information is known about its gene structure, gene regulation and gene function. The vast gene resource generated by the expressed sequence tags (ESTs) efforts was exploited to identify, using a bioinformatic approach, novel human and murine gene transcripts homologous to Drosophila mutant genes. A systematic characterization of these genes, named Drosophila-related expressed sequences (DRES), was performed including genomic mapping in human and mouse and detailed study of their expression pattern by RNA in situ hybridization experiments. Comparison between DRES genes and their putative partners in Drosophila contributes to the understanding of their function in mammals and to the discovery of their possible role in disease.

Amino Acid Sequence↗

Arabidopsis-rice: will colinearity allow gene prediction across the eudicot-monocot divide?

With the genomic sequencing of Arabidopsis nearing completion and rice sequencing very much in its infancy, a key question is whether we can exploit the Arabidopsis sequence to identify candidate genes for traits in cereal crops using a map-based approach. This requires the existence of colinearity between the Arabidopsis and cereal genomes, represented by rice, which is readily detectable using currently available resources, that is, Arabidopsis genomic sequence, rice ESTs, and genetic and physical maps. A detailed study of the colinearity remaining between two small regions of Arabidopsis chromosome 1 and rice suggests that at least in these regions of the Arabidopsis genome, conservation of gene orders with rice has been eroded to the point that it is no longer identifiable using comparative mapping. Although our analysis does not preclude that tracts of colinear gene orders may be identified using sequence comparisons or may exist in other regions of the rice and Arabidopsis genomes, it is unlikely that the extent of colinearity will be sufficient to allow map-based cross-species gene prediction and isolation. Our research also highlights the difficulties encountered in identifying orthologs using BLAST searches in incomplete sequence databases. This complicates the interpretation of comparative data among highly divergent species and limits the exploitation of Arabidopsis sequence in monocot studies.

Arabidopsis↗

Characterization and modeling of membrane proteins using sequence analysis.

The current libraries of amino acid sequences of membrane proteins are a valuable resource for the analysis of elements common to these proteins. Multiple-sequence alignment techniques and the identification of conserved features of transmembrane segments have improved the prediction of membrane protein topology. Molecular modeling in combination with structural studies or site-directed mutagenesis is proving to be a powerful link between theory and experiment. Unfortunately, the number of high-resolution structures of intrinsic membrane proteins, although increased recently, presents a restricted and perhaps biased view of membrane protein structure.

Amino Acid Sequence↗

ORFeome projects: gateway between genomics and omics.

The availability of entire genome sequences is expected to revolutionize the way in which biology and medicine are conducted for years to come. However, achieving this promise still requires significant effort in the areas of gene annotation, cloning and expression of thousands of known and heretofore unknown protein-encoding genes. Traditional technologies of manipulating genes are too cumbersome and inefficient when one is dealing with more than a few genes at a time. Entire libraries composed of all protein-encoding open reading frames (ORFs) cloned in highly flexible vectors will be needed to take full advantage of the information found in any genome sequence. The creation of such ORFeome resources using novel technologies for cloning and expressing entire proteomes constitutes an effective gateway from whole genome sequencing efforts to downstream 'omics' applications.

Animals↗

The Comprehensive Microbial Resource.

One challenge presented by large-scale genome sequencing efforts is effective display of uniform information to the scientific community. The Comprehensive Microbial Resource (CMR) contains robust annotation of all complete microbial genomes and allows for a wide variety of data retrievals. The bacterial information has been placed on the Web at http://www.tigr.org/CMR for retrieval using standard web browsing technology. Retrievals can be based on protein properties such as molecular weight or hydrophobicity, GC-content, functional role assignments and taxonomy. The CMR also has special web-based tools to allow data mining using pre-run homology searches, whole genome dot-plots, batch downloading and traversal across genomes using a variety of datatypes.

Bacteria↗

Peptide mass maps: a highly informative approach to protein identification.

A computer searching algorithm has been used to identify protein sequences in the Protein Information Resource (PIR) database with peptide mass information (mass map) obtained from proteolytic digests of proteins analyzed by microcapillary high-performance liquid chromatography electrospray ionization mass spectrometry. A theoretical analysis of the cytochrome c family demonstrates the ability to identify protein sequences in the PIR database with a high degree of accuracy using a set of six predicted tryptic peptide masses. This method was also applied to experimentally determined peptide masses for a small GTP-binding protein, a protein from pig uterus, the human sex steroid binding protein, and a thermostable DNA polymerase. The results demonstrate that a set of observed masses which is less than 50% of the total number of predicted masses can be used to identify a protein sequence in the database. For the analysis presented in this paper, a mass matching tolerance of 1 amu is used. Under these conditions, mass maps created by fast atom bombardment mass spectrometry and matrix-assisted laser desorption time-of-flight would also be applicable. In cases where multiple matches are observed or verification of the protein identification is needed, tandem mass spectrometry sequencing can be used to establish sequence similarity.

Algorithms↗

The status, quality, and expansion of the NIH full-length cDNA project: the Mammalian Gene Collection (MGC).

The National Institutes of Health's Mammalian Gene Collection (MGC) project was designed to generate and sequence a publicly accessible cDNA resource containing a complete open reading frame (ORF) for every human and mouse gene. The project initially used a random strategy to select clones from a large number of cDNA libraries from diverse tissues. Candidate clones were chosen based on 5'-EST sequences, and then fully sequenced to high accuracy and analyzed by algorithms developed for this project. Currently, more than 11,000 human and 10,000 mouse genes are represented in MGC by at least one clone with a full ORF. The random selection approach is now reaching a saturation point, and a transition to protocols targeted at the missing transcripts is now required to complete the mouse and human collections. Comparison of the sequence of the MGC clones to reference genome sequences reveals that most cDNA clones are of very high sequence quality, although it is likely that some cDNAs may carry missense variants as a consequence of experimental artifact, such as PCR, cloning, or reverse transcriptase errors. Recently, a rat cDNA component was added to the project, and ongoing frog (Xenopus) and zebrafish (Danio) cDNA projects were expanded to take advantage of the high-throughput MGC pipeline.

Animals↗

T4-induced alpha- and beta-glucosyltransferase: cloning of the genes and a comparison of their products based on sequencing data.

Bacteriophage T4 alpha- and beta-glucosyltransferases link glucosyl units to the 5-HMdC residues of its DNA. The monoglucosyl group in alpha-linkage predominates over the one in beta linkage. Having recently reported on the nucleotide sequence of gene alpha gt (1) we now determined the nucleotide sequence of gene beta gt. The genes were each cloned on a high expression vector under the control of the lambda pL promoter. After thermo-induction the proteins were isolated and purified to homogeneity. To verify that the translational starting sites and the proposed reading frames are effective in vivo the sequence of the first 31 amino acid residues from gp alpha gt and the first 30 amino acid residues from gp beta gt were determined by Edman degradation. The primary structures of the two proteins seem to have only limited structural similarities. The results are discussed comparing secondary structure predictions and homologies with other proteins from the protein sequence database of the Protein Identification Resource.

Amino Acid Sequence↗

Physical maps for genome analysis of serotype A and D strains of the fungal pathogen Cryptococcus neoformans.

The basidiomycete fungus Cryptococcus neoformans is an important opportunistic pathogen of humans that poses a significant threat to immunocompromised individuals. Isolates of C. neoformans are classified into serotypes (A, B, C, D, and AD) based on antigenic differences in the polysaccharide capsule that surrounds the fungal cells. Genomic and EST sequencing projects are underway for the serotype D strain JEC21 and the serotype A strain H99. As part of a genomics program for C. neoformans, we have constructed fingerprinted bacterial artificial chromosome (BAC) clone physical maps for strains H99 and JEC21 to support the genomic sequencing efforts and to provide an initial comparison of the two genomes. The BAC clones represented an estimated 10-fold redundant coverage of the genomes of each serotype and allowed the assembly of 20 contigs each for H99 and JEC21. We found that the genomes of the two strains are sufficiently distinct to prevent coassembly of the two maps when combined fingerprint data are used to construct contigs. Hybridization experiments placed 82 markers on the JEC21 map and 102 markers on the H99 map, enabling contigs to be linked with specific chromosomes identified by electrophoretic karyotyping. These markers revealed both extensive similarity in gene order (conservation of synteny) between JEC21 and H99 as well as examples of chromosomal rearrangements including inversions and translocations. Sequencing reads were generated from the ends of the BAC clones to allow correlation of genomic shotgun sequence data with physical map contigs. The BAC maps therefore represent a valuable resource for the generation, assembly, and finishing of the genomic sequence of both JEC21 and H99. The physical maps also serve as a link between map-based and sequence-based data, providing a powerful resource for continued genomic studies

Chromosomes, Artificial, Bacterial↗

Sample sequencing of a Salmonella typhimurium LT2 lambda library: comparison to the Escherichia coli K12 genome.

As part of the ongoing sequencing of the complete Salmonella typhimurium LT2 genome, a partly ordered set of 416 lambda clones has been developed, representing over 90% of the genome. The average insert size is 17 kb. Sequences were obtained from both ends of each clone in this set. A total of over 600 kb of sequence has been deposited in the genome survey sequence section of GenBank. This resource of clones is available from the Salmonella Genome Stock Center. A preliminary comparison with the Escherichia coli K12 genome indicates that there are likely to be many hundred insertion deletion events, encompassing more than one gene, that distinguish these genomes. Fully 30% of the S. typhimurium sequences have no close homologs in the GenBank database.

Adult↗

Leaf Ests from Stevia rebaudiana: a resource for gene discovery in diterpene synthesis.

Expressed sequence tags (ESTs) are providing a new approach to gene discovery in plant secondary metabolism. Stevia rebaudiana Bert. leaves produce high concentrations of diterpene steviol glycosides and should be a rich source of transcripts involved in diterpene synthesis. In order to create a resource for gene discovery and increase our understanding of steviol glycoside biosynthesis, we sequenced 5,548 ESTs from a S. rebaudiana leaf cDNA library. The EST collection was fully annotated based on database search results. ESTs involved in diterpene synthesis were identified using published sequences as electronic probes, by keyword searches of search results, and by differential representation. A significant portion of the ESTs were specific for standard leaf metabolic pathways; energy and primary metabolism represented 17.6% and 13.1% of total transcripts respectively. Diterpene metabolism in S. rebaudiana represented 1.1% of total transcripts. This study identified candidate genes for 70% of the known steps in the steviol glycoside pathway. One candidate, kaurene oxidase, was the 8th most abundant EST in the collection. Identification of many candidate genes specific to the I -deoxyxylulose 5-phosphate pathway suggests that the primary source of isopentenyl diphosphate, a precursor of geranylgeranyl diphosphate, is via the non-mevalonic acid pathway. The use of ESTs has greatly facilitated the identification of candidate genes and increased our understanding of diterpene metabolism.

DNA, Complementary↗

Amino acid sequences of myotoxins from Crotalus viridis concolor venom.

Myotoxins I and II were isolated from the venom of Crotalus viridis concolor. Complete sequences were derived for each reduced, alkylated toxin with data obtained by a single run on a gas phase sequencer and from fragments derived by cyanogen bromide cleavage. The results demonstrate that microheterogeneity is present in myotoxin II. The newly established sequences were compared with 3447 protein sequences in the Protein Information Resource database. The only homologous proteins found were other known myotoxins from rattlesnake venoms, namely myotoxin a, crotamine and peptide C.

Amino Acid Sequence↗

TreeGeneBrowser: phylogenetic data mining of gene sequences from public databases.

MOTIVATION: Sequence databases represent an enormous resource of phylogenetic information, but there is a lack of tools for accessing that information in order to assess the amount of evolutionary information in these databases that may be suitable for phylogenetic reconstruction and for identifying areas of the taxonomy that are under-represented for specific gene sequences. RESULTS: We have developed TreeGeneBrowser which allows inspection and evaluation of gene sequence data for phylogenetic reconstruction. This program improves the efficiency of identification of genes that may be useful for particular phylogenetic studies and identifies taxa and taxonomic branches that are under-represented in sequence databases.

Algorithms↗

MODBASE, a database of annotated comparative protein structure models, and associated resources.

MODBASE (http://salilab.org/modbase) is a relational database of annotated comparative protein structure models for all available protein sequences matched to at least one known protein structure. The models are calculated by MODPIPE, an automated modeling pipeline that relies on the MODELLER package for fold assignment, sequence-structure alignment, model building and model assessment (http:/salilab.org/modeller). MODBASE uses the MySQL relational database management system for flexible querying and CHIMERA for viewing the sequences and structures (http://www.cgl.ucsf.edu/chimera/). MODBASE is updated regularly to reflect the growth in protein sequence and structure databases, as well as improvements in the software for calculating the models. For ease of access, MODBASE is organized into different data sets. The largest data set contains 1,26,629 models for domains in 659,495 out of 1,182,126 unique protein sequences in the complete Swiss-Prot/TrEMBL database (August 25, 2003); only models based on alignments with significant similarity scores and models assessed to have the correct fold despite insignificant alignments are included. Another model data set supports target selection and structure-based annotation by the New York Structural Genomics Research Consortium; e.g. the 53 new structures produced by the consortium allowed us to characterize structurally 24,113 sequences. MODBASE also contains binding site predictions for small ligands and a set of predicted interactions between pairs of modeled sequences from the same genome. Our other resources associated with MODBASE include a comprehensive database of multiple protein structure alignments (DBALI, http://salilab.org/dbali) as well as web servers for automated comparative modeling with MODPIPE (MODWEB, http://salilab. org/modweb), modeling of loops in protein structures (MODLOOP, http://salilab.org/modloop) and predicting functional consequences of single nucleotide polymorphisms (SNPWEB, http://salilab. org/snpweb).

Amino Acid Sequence↗

Strategy for protein isoform identification from expressed sequence tags and its application to peptide mass fingerprinting.

Expressed Sequence Tags (ESTs) are an invaluable resource for protein identification and characterisation in proteomics. They allow proteins to be identified in the absence of genome sequence data. When EST sequences are used for protein identification, they are usually first processed into contigs to reduce redundancy and generate longer sequences from the overlapping ESTs. However, the process of generating contigs may accidentally group biologically meaningful isoforms together. Here we report means of discovering isoforms in EST sequences and how to use this information in the framework of protein identification and characterisation with peptide mass fingerprinting. We illustrate our strategies with examples from the dbEST database as well as protein isoforms from two-dimensional polyacrylamide gels.

Amino Acid Sequence↗