Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequencing Resource”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Sequence-ready contig for the 1.4-cM ductal carcinoma in situ loss of heterozygosity region on chromosome 8p22-p23.

We report the construction of an approximately 1.7-Mb sequence-ready YAC/BAC clone contig of 8p22-p23. This chromosomal region has been associated with frequent loss of heterozygosity (LOH) in breast, ovarian, prostate, head and neck, and liver cancer. We first constructed a meiotic linkage map for 8p to resolve previously reported conflicting map orders from the literature. The target region containing a putative tumor suppressor gene was defined by allelotyping 65 cases of sporadic ductal carcinoma in situ with 18 polymorphic markers from 8p. The minimal region of loss encompassed the interval between D8S520 and D8S261, and one tumor had loss of D8S550 only. We chose to begin physical mapping of this minimal LOH region by concentrating on the distal end, which includes D8S550. A fine-structure radiation hybrid map for the region that extends from D8S520 (distal) to D8S1759 (proximal) was prepared, followed by construction of a single, integrated YAC/BAC contig for the interval. The approximately 1730-kb contig consists of 13 YACs and 27 BACs. Fifty-four sequence-tagged sites (STSs) developed from BAC insert end-sequences and 11 expressed sequence tags were localized within the contig by STS content mapping. In addition, four unique cDNA clones from the region were isolated and fully sequenced. This integrated YAC/BAC resource provides the starting point for transcription mapping, genomic sequencing, and positional cloning of this region.

Breast Neoplasms↗

Proteome analysis of early somatic embryogenesis in Picea glauca.

Forestry is a valuable natural resource for many countries. Rapid production of large quantities of genetically improved and uniform seedlings for restocking harvested lands is a key component of sustainable forest management programs. Clonal propagation through somatic embryogenesis has the potential to meet this need in conifers and can offer the added benefit of ensuring consistent seedling quality. Although in commercial use, mass production of conifers through somatic embryogenesis is relatively new and there are numerous biological unknowns regarding this complex developmental pathway. To aid in unravelling the embryo developmental process, two-dimensional electrophoresis was employed to quantitatively assess the expression levels of proteins across four stages of somatic embryo maturation in white spruce (0, 7, 21 and 35 days post abscisic acid treatment). Forty-eight differentially expressed proteins have been identified, which display a significant change in abundance as early as day 7 of embryo development. These proteins are involved in a variety of cellular processes, many of which have not previously been associated with embryo development. The identification of these proteins was greatly assisted by the availability of a substantial expressed sequence tag (EST) resource developed for white, sitka and interior spruce. The combined use of these spruce ESTs in conjunction with GenBank accessions for other plants improved the rate of protein identification from 38% to 62% when compared with GenBank alone using automated, high-throughput techniques. This underscores the utility of EST resources in a proteomic study of any species for which a genome sequence is unavailable.

Cell Line↗

The PIR-International Protein Sequence Database.

The Protein Information Resource (PIR; http://www-nbrf.georgetown. edu/pir/) supports research on molecular evolution, functional genomics, and computational biology by maintaining a comprehensive, non-redundant, well-organized and freely available protein sequence database. Since 1988 the database has been maintained collaboratively by PIR-International, an international association of data collection centers cooperating to develop this resource during a period of explosive growth in new sequence data and new computer technologies. The PIR Protein Sequence Database entries are classified into superfamilies, families and homology domains, for which sequence alignments are available. Full-scale family classification supports comparative genomics research, aids sequence annotation, assists database organization and improves database integrity. The PIR WWW server supports direct on-line sequence similarity searches, information retrieval, and knowledge discovery by providing the Protein Sequence Database and other supplementary databases. Sequence entries are extensively cross-referenced and hypertext-linked to major nucleic acid, literature, genome, structure, sequence alignment and family databases. The weekly release of the Protein Sequence Database can be accessed through the PIR Web site. The quarterly release of the database is freely available from our anonymous FTP server and is also available on CD-ROM with the accompanying ATLAS database search program.

Amino Acid Sequence↗

MIPS bacterial genomes functional annotation benchmark dataset.

MOTIVATION: Any development of new methods for automatic functional annotation of proteins according to their sequences requires high-quality data (as benchmark) as well as tedious preparatory work to generate sequence parameters required as input data for the machine learning methods. Different program settings and incompatible protocols make a comparison of the analyzed methods difficult. RESULTS: The MIPS Bacterial Functional Annotation Benchmark dataset (MIPS-BFAB) is a new, high-quality resource comprising four bacterial genomes manually annotated according to the MIPS functional catalogue (FunCat). These resources include precalculated sequence parameters, such as sequence similarity scores, InterPro domain composition and other parameters that could be used to develop and benchmark methods for functional annotation of bacterial protein sequences. These data are provided in XML format and can be used by scientists who are not necessarily experts in genome annotation. AVAILABILITY: BFAB is available at http://mips.gsf.de/proj/bfab

Bacterial Proteins↗

ParameciumDB: a community resource that integrates the Paramecium tetraurelia genome sequence with genetic data.

ParameciumDB (http://paramecium.cgm.cnrs-gif.fr) is a new model organism database associated with the genome sequencing project of the unicellular eukaryote Paramecium tetraurelia. Built with the core components of the Generic Model Organism Database (GMOD) project, ParameciumDB currently contains the genome sequence and annotations, linked to available genetic data including the Gif Paramecium stock collection. It is thus possible to navigate between sequences and stocks via the genes and alleles. Phenotypes, of mutant strains and of knockdowns obtained by RNA interference, are captured using controlled vocabularies according to the Entity-Attribute-Value model. ParameciumDB currently supports browsing of phenotypes, alleles and stocks as well as querying of sequence features (genes, UniProt matches, InterPro domains, Gene Ontology terms) and of genetic data (phenotypes, stocks, RNA interference experiments). Forms allow submission of RNA interference data and some bioinformatics services are available. Future ParameciumDB development plans include coordination of human curation of the near 40 000 gene models by members of the research community.

Alleles↗

Generation of EST and microarray resources for functional genomic studies on chicken intestinal health.

Expressed sequenced tags (ESTs) and microarray resources have a great impact on the ability to study host response in mice and humans. Unfortunately, these resources are not yet available for domestic farm animals. The aim of this study was to provide genomic resources to study chicken intestinal health, in particular malabsorption syndrome (MAS), which affects mainly the intestine. Therefore a normalized and subtracted cDNA library containing more than 7000 clones was prepared. Randomly chosen clones were sequenced for control purposes. New ESTs were found and multiple ESTs not identified in the chicken intestine before were observed. The number of non-specific ESTs in this cDNA library was low. Based on this normalized and subtracted library a cDNA microarray was made. In a preliminary hybridization experiment with the microarray, genes were identified to be up- or downregulated in MAS infected chickens. This indicates that the generated resources are valuable tools to investigate chicken intestinal health by whole genome expression analysis approaches.

Animals↗

High-throughput computational and experimental techniques in structural genomics.

Structural genomics has as its goal the provision of structural information for all possible ORF sequences through a combination of experimental and computational approaches. The access to genome sequences and cloning resources from an ever-widening array of organisms is driving high-throughput structural studies by the New York Structural Genomics Research Consortium. In this report, we outline the progress of the Consortium in establishing its pipeline for structural genomics, and some of the experimental and bioinformatics efforts leading to structural annotation of proteins. The Consortium has established a pipeline for structural biology studies, automated modeling of ORF sequences using solved (template) structures, and a novel high-throughput approach (metallomics) to examining the metal binding to purified protein targets. The Consortium has so far produced 493 purified proteins from >1077 expression vectors. A total of 95 have resulted in crystal structures, and 81 are deposited in the Protein Data Bank (PDB). Comparative modeling of these structures has generated >40,000 structural models. We also initiated a high-throughput metal analysis of the purified proteins; this has determined that 10%-15% of the targets contain a stoichiometric structural or catalytic transition metal atom. The progress of the structural genomics centers in the U.S. and around the world suggests that the goal of providing useful structural information on most all ORF domains will be realized. This projected resource will provide structural biology information important to understanding the function of most proteins of the cell.

Algorithms↗

Structure and organization of the mitochondrial genome of the canine heartworm, Dirofilaria immitis.

This study determined the complete mitochondrial (mt) genome sequence of the canine heartworm, Dirofilaria immitis, and compared its structure, organization and other characteristics with Onchocerca volvulus and other secernentean nematodes. The D. immitis mt genome is 13814 bp in size and contains 36 of the 37 genes typical of metazoan organisms, and lacks the ATP synthetase subunit 8 gene. All of the genes are transcribed in the same direction. For the entire genome, the nucleotide contents are approximately 55% (T), approximately 19% (each for A and G) and approximately 7% (C), which is very similar to those of the protein-coding genes. In the latter genes, most (approximately 69%) third codon positions have a T, but rarely (approximately 1-9%) have an A or a C. The C content (8-12%) is higher at the first and second codon positions compared with the third position (approximately 1%). These nucleotide biases have a significant effect on the codon usage patterns and, thus, on the amino acid composition of the proteins. The mt genome organization of D. immitis is essentially the same as that of O. volvulus, but is distinctly different from other secernentean nematodes sequenced thus far. Irrespective of transpositions of transfer RNA (trn) genes and the non-coding, AT-rich region, there are 4 gene- or gene block-translocations between the mt genome of D. immitis and those of Caenorhabditis elegans, Ascaris suum and the 2 human hookworms, Ancylostoma duodenale and Necator americanus. For D. immitis, the 22 trn genes have secondary structures typical of other secernentean nematodes, and possess a TV-replacement loop instead of a TpsiC arm and loop. Like O. volvulus, the mt trnK and trnP of D. immitis use the anticodons CUU and AGG, whereas in other nematodes, UUU and UGG are employed, respectively. Also, the secondary structures of the 2 ribosomal RNA (rrn) genes are similar to the models for other nematodes. Overall, the availability of the complete D. immitis mt genome sequence provides a resource for future studies of the comparative mt genomics and of the population genetics and/or phylogeny of parasitic nematodes.

Amino Acid Sequence↗

Nematode.net: a tool for navigating sequences from parasitic and free-living nematodes.

Nematode.net (www.nematode.net) is a web- accessible resource for investigating gene sequences from nematode genomes. The database is an outgrowth of the parasitic nematode EST project at Washington University's Genome Sequencing Center (GSC), St Louis. A sister project at the University of Edinburgh and the Sanger Institute is also underway. More than 295,000 ESTs have been generated from >30 nematodes other than Caenorhabditis elegans including key parasites of humans, animals and plants. Nematode.net currently provides NemaGene EST cluster consensus sequence, enhanced online BLAST search tools, functional classifications of cluster sequences and comprehensive information concerning the ongoing generation of nematode genome data. The long-term goal of nematode.net is to provide the scientific community with the highest quality sequence information and tools for studying these diverse species.

Animals↗

What the Aspergillus genomes have told us.

The sequencing and annotation of the genomes of the first strains of Aspergillus nidulans, Aspergillus oryzae, and Aspergillus fumigatus will be seen in retrospect as a transformational event in Aspergillus biology. With this event the entire genetic composition of A. nidulans, the sexual experimental model organism of the genus Aspergillus, A. oryzae, the food biotechnology organism which is the product of centuries of cultivation, and A. fumigatus, the most common causative agent of invasive aspergillosis is now revealed to the extent that we are at present able to understand. Each genome exhibits a large set of genes common to the three as well as a much smaller set of genes unique to each. Moreover, these sequences serve as resources providing the major tool to expanding our understanding of the biology of each. Transcription profiling of A. fumigatus at high temperatures and comparative genomic hybridization between A. fumigatus and a closely related Aspergillus species provides microarray based examples of the beginning of functional analysis of the genomes of these organisms going forward from the genome sequence.

Aspergillosis↗

An organismic critique of molecular darwinism.

The molecular darwinian approach to the emergence of life treats the competition between RNA sequences for nucleotide resources as the primordial selective process in prebiotic evolution, which prescribes possible pathways for the subsequent elaboration of organizational relationships. Since success in this competition is determined by the "phenotypic" properties of RNA strands in the absence of organizational context, the genesis of biotic organization is dependent upon the establishment of co-operative, hypercyclic interactions between competing RNA sequences. The thesis of this paper is that hypercycle theory is based on unwarranted assumptions about the conditions of prebiotic evolution, and that the implications of these assumptions run counter to both empirical evidence and to the rational by which natural selection operates in evolution generally. An organismic alternative to hypercycle theory is suggested, based on the catalytic microsphere and the thermodynamics of selection.

Amino Acid Sequence↗

Generation of a transcription map of a 1 Mbase region containing the HFE gene (6p22).

A transcription map was generated of a 1 Mb interval including the HFE gene on 6p22. Thirty-seven unique cDNA fragments were characterised following their retrieval from hybridisation of immobilised YACs to primary pools of cDNAs prepared from RNA of foetal brain, human liver, foetal human liver, placenta, and CaCo2 cell line. All cDNA fragments were positioned on the physical map on the basis of presence in aligned and overlapping YACs and cosmid clones of the region. The isolated cDNAs together with established or published sequence tagged sites (STSs) and markers provided sufficient landmark density to cover approximately 90% of the 1 Mb interval with cosmid clones. The precise localisation of two known genes (NPT1 and RING finger protein) was established. A minimum of 14 additional transcription units has also been integrated. Twenty-eight cDNA fragments showed no similarity with known sequences, but 20 of these detected discrete mRNAs upon northern analysis. Their characterisation is still under investigation. Eleven new polymorphisms were also identified and localised, and the HFE genomic structure was better defined. This integrated transcription map considerably extends a recently published map of the HFE region. It will be useful for the identification of genetic defects mapping to this region and for providing template resources for genomic sequencing.

Caco-2 Cells↗

Numerous novel annotations of the human genome sequence supported by a 5'-end-enriched cDNA collection.

A collection of 90,000 human cDNA clones generated to increase the fraction of "full-length" cDNAs available was analyzed by sequence alignment on the human genome assembly. Five hundred fifty-two gene models not found in LocusLink, with coding regions of at least 300 bp, were defined by using this collection. Exon composition proposed for novel genes showed an average of 4.7 exons per gene. In 20% of the cases, at least half of the exons predicted for new genes coincided with evolutionary conserved regions defined by sequence comparisons with the pufferfish Tetraodon nigroviridis. Among this subset, CpG islands were observed at the 5' end of 75%. In-frame stop codons upstream of the initiator ATG were present in 49% of the new genes, and 16% contained a coding region comprising at least 50% of the cDNA sequence. This cDNA resource also provided candidate small protein-coding genes, usually not included in genome annotations. In addition, analysis of a sample from this cDNA collection indicates that approximately 380 gene models described in LocusLink could be extended at their 5' end by at least one new exon. Finally, this cDNA resource provided an experimental support for annotations based exclusively on predictions, thus representing a resource substantially improving the human genome annotation.

5' Untranslated Regions↗

Identification of mouse brain proteins after two-dimensional electrophoresis and electroblotting by microsequence analysis and amino acid composition analysis.

Two-dimensional electrophoretic separation and immobilization of proteins onto inert membranes for subsequent amino acid sequence and amino acid composition analysis is described as a rapid procedure for the identification or characterization of proteins from complex mixtures. This method avoids the drawbacks of classical purification and isolation methods which involve time-consuming operations with low resolution and, often, insufficient yields. Excellent overall yields of minor amounts (in the low microgram range) using this method allow for sequence determination of yet inaccessible proteins. Solubilized cell proteins of mouse brain were separated by high resolution two-dimensional electrophoresis and electroblotted onto a siliconized glass fiber membrane. The immobilized proteins were stained with Coomassie Brilliant Blue R-250, and twelve proteins spots were then submitted to both Edman degradation and amino acid analysis. Proteins were identified by comparison of the experimentally determined amino acid composition with a dataset derived from the Protein Identification Resource (PIR) protein sequence database. Eight out of twelve proteins tested were identified by amino acid analysis and confirmed by N-terminal sequence determination.

Amino Acid Sequence↗

An annotated cDNA library and microarray for large-scale gene-expression studies in the ant Solenopsis invicta.

Ants display a range of fascinating behaviors, a remarkable level of intra-species phenotypic plasticity and many other interesting characteristics. Here we present a new tool to study the molecular mechanisms underlying these traits: a tentatively annotated expressed sequence tag (EST) resource for the fire ant Solenopsis invicta. From a normalized cDNA library we obtained 21,715 ESTs, which represent 11,864 putatively different transcripts with very diverse molecular functions. All ESTs were used to construct a cDNA microarray.

Animals↗

Molecular cloning, structure, and expression of dopamine-beta-hydroxylase from bovine adrenal medulla.

Dopamine-beta-hydroxylase (DBH), the enzyme that catalyzes the conversion of dopamine to norepinephrine, remains the topic of many unanswered questions. We isolated DBH cDNA clones from a bovine adrenal medulla cDNA library in the vector lambda gt10. The longest cDNA had an open reading frame encoding an entire mature DBH 578 amino acid (64,808 dalton) polypeptide chain, though lacking a portion of the signal peptide. Additional 5' clones, obtained by the polymerase chain reaction, established the sequence of a 19 amino acid signal peptide. The mature protein sequence was 84% homologous to that of human pheochromocytoma DBH, including preservation of four potential copper ligand sites (HH or HXH) and substrate binding domains. There were no hydrophobic (putative membrane spanning) domains, other than the signal peptide. All available DBH peptide and protein sequence data can be accounted for by the cDNA-deduced 578 amino acid mature protein primary structure. Prokaryotic DBH expression yielded a 65-kilodalton DBH-immunoreactive peptide that differed from eukaryotic adrenal DBH only in N-linked, endoglycosidase F-sensitive glycosylation in the latter. Southern analysis suggested one DBH gene, whereas Northern analysis suggested a single 2.6 kbp tissue-specific DBH transcript. Comparison of the DBH primary structure with other reported sequences [Protein Identification Resource (PIR), New Atlas (NEWAT)] did not indicate that DBH is a member of any known gene family. The results suggest that a single DBH gene encodes a message specifying a single DBH polypeptide chain.

Adrenal Medulla↗

MIPDB: a relational database dedicated to MIP family proteins.

BACKGROUND INFORMATION: The MIPs (major intrinsic proteins) constitute a large family of membrane proteins that facilitate the passive transport of water and small neutral solutes across cell membranes. Since water is the most abundant molecule in all living organisms, the discovery of selective water-transporting channels called AQPs (aquaporins) has led to new knowledge on both the physiological and molecular mechanisms of membrane permeability. The MIPs are identified in Archaea, Bacteria and Eukaryota, and the rapid accumulation of new sequences in the database provides an opportunity for large-scale analysis, to identify functional and/or structural signatures or to infer evolutionary relationships. To help perform such an analysis, we have developed MIPDB (database for MIP proteins), a relational database dedicated to members of the MIP family. RESULTS: MIPDB is a motif-oriented database that integrates data on 785 MIP proteins from more than 200 organisms and contains 230 distinct sequence motifs. MIPDB proposes the classification of MIP proteins into three functional subgroups: AQPs, glycerol-uptake facilitators and aquaglyceroporins. Plant MIPs are classified into three specific subgroups according to their subcellular distribution in the plasma membrane, tonoplast or the symbiosome membrane. Some motifs of the database are highly selective and can be used to predict the transport function or subcellular localization of unknown MIP proteins. CONCLUSIONS: MIPDB offers a user-friendly and intuitive interface for a rapid and easy access to MIP resources and to sequence analysis tools. MIPDB is a web application, publicly accessible at http://idefix.univ-rennes1.fr:8080/Prot/index.html.

Amino Acid Motifs↗

VMD: a community annotation database for oomycetes and microbial genomes.

The VBI Microbial Database (VMD) is a database system designed to host a range of microbial genome sequences. At present, the database contains genome sequence and annotation data of two plant pathogens Phytophthora sojae and Phytophthora ramorum. With the completion of the draft genome sequences of these pathogens in collaboration with the DOE Joint Genome Institute (JGI), we have created this resource to make the sequences publicly available. The genome sequences (95 MB for P.sojae and 65 MB for P.ramorum) were annotated with approximately 19,000 and approximately 16,000 gene models, respectively. We used two different statistical methods to validate these gene models, Fickett's and a log-likelihood method. Functional annotation of the gene models is based on results from BlastX and InterProScan screens. From the InterProScan results, we could assign putative functions to 17,694 genes in P.sojae and 14,700 genes in P.ramorum. We created an easy-to-use genome browser to view the genome sequence data, which opens to detailed annotation pages for each gene model. A community annotation interface is available for registered community members to add or edit annotations. There are approximately 1600 gene models for P.sojae and approximately 700 models for P.ramorum that have already been manually curated. A toolkit is provided as an additional resource for users to perform a variety of sequence analysis jobs. The database is publicly available at http://phytophthora.vbi.vt.edu/.

Databases, Nucleic Acid↗