Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequencing Resource”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

Nucleic acid sequence database VI: Retroviral oncogenes and cellular proto-oncogenes.

The databases of the Protein Identification Resource at the National Biomedical Research Foundation (NBRF) contain nucleic acid and protein sequences from 18 retroviral oncogenes (v-onc) and 8 cellular proto-oncogenes (c-onc). Comparison of the sequences between the v-onc and c-onc genes reveals: (i) The c-src, c-abl, c-mos, c-fos, c-ras, c-myb, c-myc, and c-sis genes contain coding regions that are highly conserved in the respective v-onc genes with a small number of base changes. (ii) There are more transitions than transversions. (iii) Some of these base changes are silent mutations and others generate amino acid substitutions in the viral proteins. The causes of these base changes in the coding sequences and the significance to oncogenic transformation of the amino acid substitutions in the viral proteins remain to be determined.

Animals↗

An integrated genetic map and a new set of simple sequence repeat markers for pearl millet, Pennisetum glaucum.

Over the past 10 years, resources have been established for the genetic analysis of pearl millet, Pennisetum glaucum (L.) R. Br., an important staple crop of the semi-arid regions of India and Africa. Among these resources are detailed genetic maps containing both homologous and heterologous restriction fragment length polymorphism (RFLP) markers, and simple sequence repeats (SSRs). Genetic maps produced in four different crosses have been integrated to develop a consensus map of 353 RFLP and 65 SSR markers. Some 85% of the markers are clustered and occupy less than a third of the total map length. This phenomenon is independent of the cross. Our data suggest that extreme localization of recombination toward the chromosome ends, resulting in gaps on the genetic map of 30 cM or more in the distal regions, is typical for pearl millet. The unequal distribution of recombination has consequences for the transfer of genes controlling important agronomic traits from donor to elite pearl millet germplasm. The paper also describes the generation of 44 SSR markers from a (CA)n-enriched small-insert genomic library. Previously, pearl millet SSRs had been generated from BAC clones, and the relative merits of both methodologies are discussed.

Africa↗

Identification of polymorphisms in the human Reprimo gene using public EST data.

The human Reprimo gene is a recently identified cytoplasmic protein, which plays an important role in the regulation of p53-dependent G2 arrest of the cell cycle. Genetic variations in the Reprimo gene that may influence enzyme activity can be of both biological and epidemiological significance. The human expressed sequence tag (EST) database is a wealth of resources, which can be used to rapidly screen for potential polymorphisms in proteins of physiological interest. On the basis of the alignment of human EST sequences, we identified two candidate polymorphisms at nucleotides 824 and 839 in the 3'-untranslated region of the Reprimo gene. The presence of these polymorphisms was confirmed in a Caucasian population (n=82) by the use of the allele specific polymerase chain reaction (PCR). The rare allele frequency at position 824 (38.4%) is much higher than rare allele frequency at position 839 (3.7%). Our results suggest that the human EST data may serve as a valuable source for the rapid identification of genetic variation.

3' Untranslated Regions↗

Whole genome sequence of a superbug-Escherichia coli strain KAB-AI-497 isolated from the vagina of a 20 year old pregnant woman with premature rupture of membrane (PROM) in a resource limited setting, Kabale Regional Referral Hospital, in Uganda.

OBJECTIVES: The objective of the study is to sequence the whole genome of multidrug resistant E. coli strain KAB-AI-497 that causes bacterial vaginosis and implicated in premature rupture of membrane in pregnant woman. DATA DESCRIPTION: The DNA of the E. coli strain KAB-AI-497 was extracted using the MagAttract HMW DNA Kit, and the extracted DNA was sequenced using an MGI DNBSEQ G99ARS platform. FastQC was used to perform quality control analysis and the reads were trimmed by Trimmomatic. De novo genome assembly was performed by SPAdes and it resulted to a draft assembled genome that has 5.1 Mb genome size, 153 contigs, and 50.5% GC content. Quality analysis of the assembled genome revealed it has 98.46% completeness and 0.97% contamination. The closest E. coli strain to this strain KAB-AI-497 in terms of similarity was Escherichia coli SMS-3-5 with an average nucleotide identity of 98.43% and genome coverage of 86.19%, which confirmed the species level identity of the strain. The assembled genome was annotated using the NCBI Prokaryotic Genome Annotation Pipeline which identified 4,726 protein coding genes in the strain genome. Furthermore, the annotation revealed the genome has resistant genes responsible for resistance against many antibiotic classes such as tetracycline, fluoroquinolone, and penicillin.

Escherichia coli↗

The RESID Database of Protein Modifications: 2003 developments.

The RESID Database is a comprehensive collection of annotations and structures for protein pre-, co- and post-translational modifications including amino-terminal, carboxyl-terminal and peptide chain cross-link modifications. The RESID Database includes: systematic and alternate names, atomic formulas and masses, enzyme activities generating the modifications, keywords, literature citations, Gene Ontology cross-references, Protein Information Resource (PIR) and SWISS-PROT protein sequence database feature table annotations, structure diagrams and molecular models. This database is freely accessible on the Internet through the European Bioinformatics Institute at http://srs.ebi.ac.uk/srs6bin/cgi-bin/wgetz?-page+LibInfo+-lib+RESID, through the National Cancer Institute - Frederick Advanced Biomedical Computing Center at http://www.ncifcrf.gov/RESID, or through the Protein Information Resource at http://pir.georgetown.edu/pirwww/dbinfo/resid.html.

Animals↗

The Centre for Modeling Human Disease Gene Trap resource.

Gene trap mutagenesis of mouse embryonic stem cells generates random loss-of-function mutations, which can be identified by a sequence tag and can often report the endogenous expression of the mutated gene. The Centre for Modeling Human Disease is performing expression- and sequence-based screens of gene trap insertions to generate new mouse mutations as a resource for the scientific community. The gene trap insertions are screened using multiplexed in vitro differentiation and induction assays, and sequence tags are generated to complement expression profiles. Researchers may search for insertions in genes expressed in target cell lineages, under specific in vitro conditions, or based upon sequence identity via an online searchable database (http://www.cmhd.ca/sub/genetrap.asp). The clones are available as a resource to researchers worldwide to help to functionally annotate the mammalian genome and will serve as a source to test candidate loci identified by phenotype-driven mutagenesis screens.

Animals↗

Cataloging transcription factor and major signaling molecule genes for functional genomic studies in Ciona intestinalis.

The ascidian Ciona intestinalis provides an excellent experimental system for functional genomic studies because (1) its genome has been sequenced, (2) the transcription factor genes and genes for major signal transduction molecules have been extensively screened and annotated on a genome-wide scale using the molecular phylogenetical method, and (3) their embryonic expression profiles have been almost completely determined. However, the entire genetic structure, including the 5' and 3' untranslated regions and the protein-coding regions, of most gene models used in these prior studies is not always supported by cDNA evidence, and thus, these gene models are potentially imprecise. To facilitate functional genomic studies based on precise gene structures, our present study determined 406 cDNA sequences for 357 transcription factor genes and 112 cDNA sequences for 107 signal transduction molecule genes, greatly improving the previous gene models and revealing transcript variants for 44 genes. Considering these data alongside those of previously characterized genes deposited in the DNA Data Bank of Japan/European Molecular Biology Laboratory/GENBANK databases, 95.6% of the catalogued transcription factor genes (373/390) and 98.3% of the catalogued signal transduction molecule genes (117/119) have now been verified by cDNA sequences. Thus, the present study greatly improves the resources available for functional genomic studies in C. intestinalis.

Animals↗

An initial strategy for the systematic identification of functional elements in the human genome by low-redundancy comparative sequencing.

With the recent completion of a high-quality sequence of the human genome, the challenge is now to understand the functional elements that it encodes. Comparative genomic analysis offers a powerful approach for finding such elements by identifying sequences that have been highly conserved during evolution. Here, we propose an initial strategy for detecting such regions by generating low-redundancy sequence from a collection of 16 eutherian mammals, beyond the 7 for which genome sequence data are already available. We show that such sequence can be accurately aligned to the human genome and used to identify most of the highly conserved regions. Although not a long-term substitute for generating high-quality genomic sequences from many mammalian species, this strategy represents a practical initial approach for rapidly annotating the most evolutionarily conserved sequences in the human genome, providing a key resource for the systematic study of human genome function.

Animals↗

REGANOR: a gene prediction server for prokaryotic genomes and a database of high quality gene predictions for prokaryotes.

UNLABELLED: With >1,000 prokaryotic genome sequencing projects ongoing or already finished, comprehensive comparative analysis of the gene content of these genomes has become viable. To allow for a meaningful comparative analysis, gene prediction of the various genomes should be as accurate as possible. It is clear that improving the state of genome annotation requires automated gene identification methods to cope with the influence of artifacts, such as genomic GC content. There is currently still room for improvement in the state of annotations. We present a web server and a database of high-quality gene predictions. The web server is a resource for gene identification in prokaryote genome sequences. It implements our previously described, accurate gene finding method REGANOR. We also provide novel gene predictions for 241 complete, or almost complete, prokaryotic genomes. We demonstrate how this resource can easily be utilised to identify promising candidates for currently missing genes from genome annotations with several examples. All data sets are available online. AVAILABILITY: The gene finding server is accessible via https://www.cebitec.uni-bielefeld.de/groups/brf/software/reganor/cgi-bin/reganor_upload.cgi. The server software is available with the GenDB genome annotation system (version 2.2.1 onwards) under the GNU general public license. The software can be downloaded from https://sourceforge.net/projects/gendb/. More information on installing GenDB and REGANOR and the system requirements can be found on the GenDB project page http://www.cebitec.uni-bielefeld.de/groups/brf/software/wiki/GenDBWiki/AdministratorDocumentation/GenDBInstallation

Chromosome Mapping↗

How physicians can improve patients' participation and maintenance in self-care.

A protocol for the stepped education and support of patients is derived from the cumulative experience of more than 200 clinical trials of patient education and behavioral change interventions. The recommended procedure entails assessing a patient's educational needs by asking a sequence of "diagnostic" questions to assure patient motivation, skill and resources and to reinforce adherence to the prescribed medical regimen or life-style modifications. The sequence of questions and interventions is also designed to minimize a physician's time commitment and to maximize the medical benefit to the patient.

Humans↗

Integrated STS/YAC physical, genetic, and transcript map of human Xq21.3 to q23/q24 (DXS1203-DXS1059).

A map has been assembled that extends from the XY homology region in Xq21.3 to proximal Xq24, approximately 20 Mb, formatted with 200 STSs that include 25 dinucleotide repeat polymorphic markers and more than 80 expressed sequences including 30 genes. New genes HTRP5, CAPN6, STPK, 14-3-3PKR, and CALM1 and previously known genes including BTK, DDP, GLA, PLP, COL4A5, COL4A6, PAK3, and DCX are localized; candidate loci for other disorders for which genes have not yet been identified, including DFN-2, POF, megalocornea, and syndromic and nonsyndromic mental retardation, are also mapped in the region. The telomeric end of the contig overlaps a yeast artificial chromosome (YAC) contig from Xq24-q26 and with other previously published contigs provides complete sequence-tagged site (STS)/YAC-based coverage of the long arm of the X chromosome. The order of published landmark loci in genetic and radiation hybrid maps is in general agreement. Combined with high-density STS landmarks, the multiple YAC clone coverage and integrated genetic, radiation hybrid, and transcript map provide resources to further disease gene searches and sequencing.

Chromosome Mapping↗

Optimized strategies for sequence-tagged-site selection in genome mapping.

The physical mapping of complex genomes is based on the construction of a genomic library and the determination of the overlaps between the inserts of the mapping clones in order to generate an ordered, cloned representation of nearly all the sequences present in the target genome. Evaluation of the relative efficiency of experimental procedures used to accomplish this goal must minimally include a comparison of the fraction of the genome covered by the ordered arrays (or "contigs"), the average size of the contigs, and the cost, in terms of time and resources, required to generate the map. Sequence-tagged-site (STS) content mapping is one strategy that has been proposed and is being utilized for this type of experiment. This paper describes three STS selection schemes and presents computer simulations of contig-building experiments based on these procedures. The results of these simulations suggest that a nonrandom STS strategy that uses paired probes requires one-third to one-fourth as many STS assays as are required in random and nonpaired approaches, and also results in a map that has both greater genome coverage and a larger average contig size. This strategy promises to reduce the time and cost required to build a high-quality physical map.

Base Sequence↗

The Gene Resource Locator: gene locus maps for transcriptome analysis.

Since the advent of the draft human genome sequence there has been growing interest in transcriptome analysis based on genomic data. The Gene Resource Locator (GRL) assembles gene maps that include information on gene-expression patterns, cis-elements in regulatory regions and alternatively spliced transcripts. The database was constructed using customized software, and currently contains 2.2 million alignments (exon-intron structures). The alignments have been annotated and integrated into a system that encompasses approximately 90 000 EST loci sharing common exons, 8091 alternatively spliced transcript groups, 10 801 expression-profile groups, 8066 candidate regulatory regions in full-length cDNAs, and 1 million SNP loci. We have used Flash technology to build a dynamic web viewer that facilitates browsing through the millions of alignments. All of the information is available through the World Wide Web at the Gene Resource Locator web site (http://grl.gi.k.u-tokyo.ac.jp).

Alternative Splicing↗

Refined annotation of the Arabidopsis genome by complete expressed sequence tag mapping.

Expressed sequence tags (ESTs) currently encompass more entries in the public databases than any other form of sequence data. Thus, EST data sets provide a vast resource for gene identification and expression profiling. We have mapped the complete set of 176,915 publicly available Arabidopsis EST sequences onto the Arabidopsis genome using GeneSeqer, a spliced alignment program incorporating sequence similarity and splice site scoring. About 96% of the available ESTs could be properly aligned with a genomic locus, with the remaining ESTs deriving from organelle genomes and non-Arabidopsis sources or displaying insufficient sequence quality for alignment. The mapping provides verified sets of EST clusters for evaluation of EST clustering programs. Analysis of the spliced alignments suggests corrections to current gene structure annotation and provides examples of alternative and non-canonical pre-mRNA splicing. All results of this study were parsed into a database and are accessible via a flexible Web interface at http://www.plantgdb.org/AtGDB/.

Alternative Splicing↗

Mapping human telomere regions with YAC and P1 clones: chromosome-specific markers for 27 telomeres including 149 STSs and 24 polymorphisms for 14 proterminal regions.

A YAC library enriched for telomere clones was constructed and screened for the human telomere-specific repeat sequence (TTAGGG). Altogether 196 TYAC library clones were studied: 189 new TYAC clones were isolated, 149 STSs were developed for 132 different TY-ACs, and 39 P1 clones were identified using 19 STSs from 16 of the TYACs. A combination of mapping methods including fluorescence in situ hybridization, somatic cell hybrid panels, clamped homogeneous electric fields, meiotic linkage, and BLASTN sequence analysis was utilized to characterize the resource. Forty-five of the TYACs map to 31 specific telomere regions. Twenty-four linkage markers were developed and mapped within 14 proterminal regions (12 telomeres and 2 terminal bands). The polymorphic markers include 12 microsatellites for 10 telomeres (1q, 2p, 6q, 7q, 10p, 10q, 13q, 14q, 18p, 22q) and the terminal bands of 11q and 12p. Twelve RFLP markers were identified and meiotically mapped to the telomeres of 2q, 7q, 8p, and 14q. Chromosome-specific STSs for 27 telomeres were identified from the 196 TYACs. More than 30,000 nucleotides derived from the TYAC vector-insert junction regions or from regions flanking TYAC microsatellites were compared to reported sequences using BLASTN. In addition to identifying homology with previously reported telomere sequences and human repeat elements, gene sequences and a number of ESTs were found to be highly homologous to the TYAC sequences. These genes include human coagulation factor V (F5), Weel protein tyrosine kinase (WEE1), neurotropic protein tyrosine kinase type 2 (NTRE2), glutathione S-transferase (GST1), and beta tubulin (TUBB). The TYAC/P1 resource, derivative STSs, and polymorphisms constitute an enabling resource to further studies of telomere structure and function and a means for physical and genetic map integration and closure.

Animals↗

The InterPro database, an integrated documentation resource for protein families, domains and functional sites.

Signature databases are vital tools for identifying distant relationships in novel sequences and hence for inferring protein function. InterPro is an integrated documentation resource for protein families, domains and functional sites, which amalgamates the efforts of the PROSITE, PRINTS, Pfam and ProDom database projects. Each InterPro entry includes a functional description, annotation, literature references and links back to the relevant member database(s). Release 2.0 of InterPro (October 2000) contains over 3000 entries, representing families, domains, repeats and sites of post-translational modification encoded by a total of 6804 different regular expressions, profiles, fingerprints and Hidden Markov Models. Each InterPro entry lists all the matches against SWISS-PROT and TrEMBL (more than 1,000,000 hits from 462,500 proteins in SWISS-PROT and TrEMBL). The database is accessible for text- and sequence-based searches at http://www.ebi.ac.uk/interpro/. Questions can be emailed to interhelp@ebi.ac.uk.

Databases, Factual↗