Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequencing Resource”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

Plant protein annotation in the UniProt Knowledgebase.

The Swiss-Prot, TrEMBL, Protein Information Resource (PIR), and DNA Data Bank of Japan (DDBJ) protein database activities have united to form the Universal Protein Resource (UniProt) Consortium. UniProt presents three database layers: the UniProt Archive, the UniProt Knowledgebase (UniProtKB), and the UniProt Reference Clusters. The UniProtKB consists of two sections: UniProtKB/Swiss-Prot (fully manually curated entries) and UniProtKB/TrEMBL (automated annotation, classification and extensive cross-references). New releases are published fortnightly. A specific Plant Proteome Annotation Program (http://www.expasy.org/sprot/ppap/) was initiated to cope with the increasing amount of data produced by the complete sequencing of plant genomes. Through UniProt, our aim is to provide the scientific community with a single, centralized, authoritative resource for protein sequences and functional information that will allow the plant community to fully explore and utilize the wealth of information available for both plant and non-plant model organisms.

Amino Acid Sequence↗

Generation of EST and cDNA microarray resources for the study of bovine immunobiology.

Recent developments in expressed sequence tag (EST) and cDNA microarray technology have had a dramatic impact on the ability of scientists to study the responses of thousands of genes to external stimuli, such as infection, nutrient flux, and stress. To date however, these studies have largely been limited to human and rodent systems. Despite the tremendous potential benefit of EST and cDNA microarray technology to studies of complex problems in domestic animal species, a lack of integrated resources has precluded application of these technologies to domestic species. To address this problem, the Center for Animal Functional Genomics (CAFG) at Michigan State University has developed a normalized bovine total leukocyte (BOTL) cDNA library, generated EST clones from this library, and printed cDNA microarrays suitable for studying bovine immunobiology. Our data revealed that the normalization procedure successfully reduced highly abundant cDNA species while enhancing the relative percentage of clones representing rare transcripts. To date, a total of 932 EST sequences have been generated from this library (BOTL) and the sequence information plus BLAST results made available through a web-accessible database (http://gowhite.ans.msu.edu). Cluster analysis of the data indicates that a total of 842 unique cDNAs are present in this collection, reflecting a low redundancy rate of 9.7%. For creation of first generation cDNA microarrays, inserts from 720 unique clones in this library were amplified and microarrays were produced by spotting each insert or amplicon 3 times on glass slides in a 48-patch arrangement with 64 total spots (including blanks and positive controls) per patch. To test our BOTL microarray, we compared gene expression patterns of concanavalin A stimulated and unstimulated peripheral blood mononuclear cells (PBMCs). In total, hybridization signals on over 90 amplicons showed upregulation (> 3x) in response to Con A stimulation, relative to unstimulated cells. A second experiment with PBMCs from a different group of animals was performed to test reproducibility of microarray results. There was a high correlation between the 2 experiments (r = 0.72, P < 0.001). Resources described in this publication offer a highly efficient and integrated system to study gene expression changes in bovine leukocytes.

Animals↗

Development of a 1.4-Mb BAC/PAC contig and physical map within the critical region for complete X-linked congenital stationary night blindness in Xp11.4.

A physical map internal to the markers DXS1368 and DXS228 was developed for the p11.4 region of the human X chromosome. Twenty-four BACs and 10 PACs with an average insert size of 149 kb were aligned to form a contig across an estimated 1.4 Mb of DNA. This contig, which has on average fourfold clone coverage, was assembled by STS and EST content analysis using 46 markers, including 8 ESTs, two retinally expressed genes, and 22 new STSs developed from BAC- and PAC-derived DNA sequence. The average intermarker distance was 30 kb. This physical map provides resources for high-resolution mapping as well as suitable clones for large-scale sequencing efforts in Xp11.4, a region known to contain the gene for complete X-linked congenital stationary night blindness.

Bacteriophages↗

Gene identification and analysis of transcripts differentially regulated in fracture healing by EST sequencing in the domestic sheep.

BACKGROUND: The sheep is an important model animal for testing novel fracture treatments and other medical applications. Despite these medical uses and the well known economic and cultural importance of the sheep, relatively little research has been performed into sheep genetics, and DNA sequences are available for only a small number of sheep genes. RESULTS: In this work we have sequenced over 47 thousand expressed sequence tags (ESTs) from libraries developed from healing bone in a sheep model of fracture healing. These ESTs were clustered with the previously available 10 thousand sheep ESTs to a total of 19087 contigs with an average length of 603 nucleotides. We used the newly identified sequences to develop RT-PCR assays for 78 sheep genes and measured differential expression during the course of fracture healing between days 7 and 42 postfracture. All genes showed significant shifts at one or more time points. 23 of the genes were differentially expressed between postfracture days 7 and 10, which could reflect an important role for these genes for the initiation of osteogenesis. CONCLUSION: The sequences we have identified in this work are a valuable resource for future studies on musculoskeletal healing and regeneration using sheep and represent an important head-start for genomic sequencing projects for Ovis aries, with partial or complete sequences being made available for over 5,800 previously unsequenced sheep genes.

Animals↗

Cytochrome P-450terp. Isolation and purification of the protein and cloning and sequencing of its operon.

Cytochromes P-450 are extremely important in the oxidative metabolism of a variety of endogenous and exogenous compounds in pro- and eukaryotic organisms. Progress in understanding the structure and mechanism of action of this superfamily of enzymes has been hampered by the properties of the eukaryotic enzymes and the availability of only one well-characterized prokaryotic enzyme as a model. We report here the isolation of a Pseudomonas species which will utilize a monoterpene natural product, alpha-terpineol, as its sole source of carbon and energy. Approximately 1% of the soluble protein in the cell-free extract is a novel cytochrome P-450 (P-450terp). This enzyme and its associated iron sulfur protein electron carrier (terpredoxin) have been purified to homogeneity and their NH2-terminal amino acid sequences determined. The amino acid sequences of six tryptic peptide fragments of cytochrome P-450terp have also been determined. This sequence information was used to clone the gene encoding cytochrome P-450terp. Three clones representing approximately 8 kilobase pairs of unique sequences were selected and sequenced. Five non-overlapping open reading frames (ORFs) were found in the sequences, and the translated sequences were used to search the Protein Identification Resource for comparable proteins. The ORFs were identified as: 1) an alcohol dehydrogenase, 2) an aldehyde dehydrogenase, 3) cytochrome P-450terp, 4) terpredoxin reductase, and 5) terpredoxin. The identification of both the cytochrome P-450terp and terpredoxin DNA sequence was confirmed by the presence of each of the corresponding amino acid sequences found in the purified proteins. The five ORFs were bounded on both the 5' and 3' ends by consensus factor-independent terminator sequences. A consensus promoter sequence was found immediately 5' to the first ORF. These results indicate that we have sequenced the complete terp operon. Comparison of the amino acid sequence of cytochrome P-450terp to that of all other cytochromes P-450 has shown that it is the first member of the gene family CYP108. Preliminary characterization of the chemical and physical properties and the preparation of crystals of this new cytochrome P-450, suitable for x-ray diffraction analysis, indicate that it will be useful in comparison studies with other members of this class of proteins.

Alcohol Dehydrogenase↗

Comparative genome assembly.

One of the most complex and computationally intensive tasks of genome sequence analysis is genome assembly. Even today, few centres have the resources, in both software and hardware, to assemble a genome from the thousands or millions of individual sequences generated in a whole-genome shotgun sequencing project. With the rapid growth in the number of sequenced genomes has come an increase in the number of organisms for which two or more closely related species have been sequenced. This has created the possibility of building a comparative genome assembly algorithm, which can assemble a newly sequenced genome by mapping it onto a reference genome. We describe here a novel algorithm for comparative genome assembly that can accurately assemble a typical bacterial genome in less than four minutes on a standard desktop computer. The software is available as part of the open-source AMOS project.

Algorithms↗

Use of a proteome strategy for tagging proteins present at the plasma membrane.

A plasma membrane (PM) fraction was purified from Arabidopsis thaliana using a standard procedure and analyzed by two-dimensional (2D) gel electrophoresis. The proteins were classified according to their relative abundance in PM or cell membrane supernatant fractions. Eighty-two of the 700 spots detected on the PM 2D gels were microsequenced. More than half showed sequence similarity to proteins of known function. Of these, all the spots in the PM-specific and PM-enriched fractions, together with half of the spots with similar abundance in PM fraction and supernatant, have previously been found at the PM, supporting the validity of this approach. Extrapolation from this analysis indicates that (i) approximately 550 polypeptides found at the PM could be resolved on 2D gels; (ii) that numerous proteins with multiple locations are found at the PM; and (iii) that approximately 80% of PM-specific spots correspond to proteins with unknown function. Among the later, half are represented by ESTs or cDNAs in databases. In this way, several unknown gene products were potentially localized to the PM. These data are discussed with respect to the efficiency of organelle proteome approaches to link systematically genomic data to genome expression. It is concluded that generalized proteomes can constitute a powerful resource, with future completion of Arabidopsis genome sequencing, for genome-wide exploration of plant function.

Amino Acid Sequence↗

A 2.8-Mb clone contig of the multiple endocrine neoplasia type 1 (MEN1) region at 11q13.

Multiple endocrine neoplasia type 1 (MEN1) is an autosomal dominant disorder that results in parathyroid, anterior pituitary, and pancreatic and duodenal endocrine tumors in affected individuals. The MEN1 locus is tightly linked to the marker PYGM on chromosome 11q13, and linkage analysis has placed the MEN1 gene within a 2-Mb interval flanked by D11S1883 and D11S449. As a step toward cloning the MEN1 gene, we have constructed a 2.8-Mb clone contig consisting of YAC and bacterial clones (PAC, BAC, and P1) for the D11S480 to D11S913 region. The bacterial clones alone represent nearly all of the 2.8-Mb contig. The contig was assembled based on a high-density STS-content analysis of 79 genomic clones (YAC, PAC, BAC, and P1) with 118 STSs. The STSs included 22 polymorphic markers and 20 transcripts, with the remainder primarily derived from the end sequences of the genomic clones. An independent cosmid contig for the 1-Mb PYGM-SEA region was also generated. Support for correctness of the 2.8-Mb contig map comes from an independent ordering of the clones by fiber-FISH. This sequence-ready contig will be a useful resource for positional cloning of MEN1 and other disease genes whose loci fall within this region.

Chromosomes, Artificial, Yeast↗

EST databases as a source for molecular markers: lessons from Helianthus.

Expressed sequence tag (EST) databases represent a potentially valuable resource for the development of molecular markers for use in evolutionary studies. Because EST-derived markers come from transcribed regions of the genome, they are likely to be conserved across a broader taxonomic range than are other sorts of markers. This paper describes a case study in which the publicly available cultivated sunflower (Helianthus annuus) EST database was used to develop simple sequence repeat (SSR) markers for use in the genetic analysis of a rare sunflower species, Helianthus verticillatus, as well as the more widespread Helianthus angustifolius. EST-derived SSRs were found to be more than 3 times as transferable across species as compared with anonymous SSRs (73% vs. 21%, respectively). Moreover, EST-SSRs whose primers were located within protein-coding sequence were more readily transferable than those derived from untranslated regions, and the former loci were no less variable than the latter. The utility of existing EST databases as a means for facilitating population genetic analyses in plants was further explored by cross-referencing publicly available EST resources against available lists of rare or invasive flowering plant taxa. This survey revealed that more than one-third of all plant-derived EST collections of sufficient size could conceivably serve as a source of EST-SSRs for the analysis of rare, endangered, or invasive plant species worldwide.

Asteraceae↗

The Rice PIPELINE: a unification tool for plant functional genomics.

The Rice Genome Research Project in Japan performs genome sequencing and comprehensive expression profiling, constructs genetic and physical maps, collects full-length cDNAs and generates mutant lines, all aimed at improving the breeding of the rice plant as a food source. The National Institute of Agrobiological Sciences in Tsukuba, Japan, has accumulated numerous rice biological resources and has already successfully produced a high-quality genome sequence, a high-density genetic map with 3000 markers, 30,000 full-length cDNAs, over 700 expression profiles with a 9000 cDNA microarray and 15,000 flanking sequences with Tos17 insertions in about 3765 mutant lines from about 50,000 transposon insertion lines. These resources are available in the public domain. A new unification tool for functional genomics, called Rice PIPELINE, has also been developed for the dynamic collection and compilation of genomics data (genome sequences, full-length cDNAs, gene expression profiles, mutant lines, cis elements) from various databases. The mission of Rice PIPELINE is to provide a unique scientific resource that pools publicly available rice genomic data for search by clone sequence, clone name, GenBank accession number, or keyword. The web-based form of Rice PIPELINE is available at http://cdna01.dna.affrc.go.jp/PIPE/.

Computational Biology↗

Identification of protein A-binding components in Spisula oocytes.

Components involved in sustaining meiosis arrest of oocytes were determined. Proteins that bind to protein A from meiosis-arrested and 5-HT-matured Spisula oocytes were analyzed by sodium dodecyl sulfate polyacrylamide gel electrophoresis. Meiosis-arrested oocytes contained three doublets of proteins with estimated Mrs of 43 and 45, 38 and 40, and 21 and 23 kDa. In 5 HT-matured oocytes the 21 and 23 and 38 and 40 kDa proteins were retained; whereas the 43 and 45 kDa proteins were absent. The protein A-bound proteins did not interact with antibodies against the various subclasses of human, mouse, rat and rabbit IgG or human Fc fragment. The amino acid sequence of the N-terminus of the 43 kDa protein was determined to be NH2-VLRIGSGMXDT. Comparison of this sequence with existing database at Protein Identification Resource (R 32.0), GenBank (R 72.0), SWISS-PROT (R 22.0), and EMBL (R 31.0) showed no homology with any reported protein. The protein A-bound components from meiosis-arrested oocytes were incubated in vitro with [gamma-32P]ATP. Only the 68 kDa protein was radiophosphorylated. This protein was not detected in 5-HT-matured oocytes. The disappearance of the 43, 45, and 68 kDa proteins in 5-HT-matured oocytes suggests that these components may be involved in sustaining meiosis meiosis. A unique property of these proteins is that they interact with protein A and are distinctly different from immunoglobulin.

Adenosine Triphosphate↗

Transition of Staphylococcus aureus tetracycline resistance plasmid pT181 from independent multicopy replicon to predominantly integrated chromosomal element over 65 years.

Mobile genetic elements (MGEs), including plasmids, phages and genome islands, are major sources of bacterial genetic diversity. The small plasmid pT181 confers tetracycline resistance in bacterial pathogen Staphylococcus aureus via an efflux pump, TetK. pT181 was one of the earliest sequenced S. aureus plasmids, and has been isolated in both clinical and livestock-associated strains for decades, both as an independent replicon and integrated in the chromosome as part of staphylococcal cassette chromosome mec (SCCmec). Bacterial genome analysis tools and high-quality sequences with metadata are publicly available, but these resources remain underleveraged for examining historical data, especially when studying the spread of MGEs across a species and over time. Using publicly available reads and metadata, we explored the evolution of pT181 over almost seven decades of samples to identify temporal trends in sequence evolution, copy number changes, and spread across S. aureus and beyond. pT181 was prevalent across S. aureus (found in 9.5% of 83,366 genomes tested), with a conserved sequence outside of three hypervariable regions. The history of pT181 since 1954 is characterized by spread across strains, significant variation in plasmid copy number of the independent replicon, and increasing frequency of integration of the plasmid into the S. aureus chromosome. We have identified multiple chromosomal integration locations of the plasmid, including outside of the previously characterized SCCmec. We find that pT181 has been transferred across staphylococcaceae and into a Gram-negative species. The repeated integration of pT181 into the chromosome may indicate co-evolution of the plasmid and the host, potentially to facilitate increased antibiotic resistance.

Journal Article↗

PhosphoSite: A bioinformatics resource dedicated to physiological protein phosphorylation.

PhosphoSite is a curated, web-based bioinformatics resource dedicated to physiologic sites of protein phosphorylation in human and mouse. PhosphoSite is populated with information derived from published literature as well as high-throughput discovery programs. PhosphoSite provides information about the phosphorylated residue and its surrounding sequence, orthologous sites in other species, location of the site within known domains and motifs, and relevant literature references. Links are also provided to a number of external resources for protein sequences, structure, post-translational modifications and signaling pathways, as well as sources of phospho-specific antibodies and probes. As the amount of information in the underlying knowledgebase expands, users will be able to systematically search for the kinases, phosphatases, ligands, treatments, and receptors that have been shown to regulate the phosphorylation status of the sites, and pathways in which the phosphorylation sites function. As it develops into a comprehensive resource of known in vivo phosphorylation sites, we expect that PhosphoSite will be a valuable tool for researchers seeking to understand the role of intracellular signaling pathways in a wide variety of biological processes.

Animals↗

Tools and resources for identifying protein families, domains and motifs.

With the large influx of raw sequence data from genome sequencing projects, there is a need for reliable automatic methods for protein sequence analysis and classification. The most useful tools use various methods for identifying motifs or domains found in previously characterized protein families. This article reviews the tools and resources available on the web for identifying signatures within proteins and discusses how they may be used in the analysis of new or unknown protein sequences.

Amino Acid Motifs↗

Nanopore Sequencing for Chikungunya Virus: Principles and Application.

Nanopore sequencing is transforming viral genomics through real-time, portable, long-read analysis of RNA and DNA. Unlike traditional short-read platforms, it detects nucleotide sequences by measuring ionic current changes as nucleic acids pass through nanoscale pores, enabling direct single-molecule sequencing and base modification detection. Its simplicity, flexibility, and capacity for ultra-long reads make it ideal for resolving complex genomic regions, structural variants, and full viral genomes. These advantages have accelerated its use in pathogen surveillance and outbreak response, especially in resource-limited settings. For chikungunya virus (CHIKV), nanopore sequencing allows rapid, culture-independent recovery of complete genomes from clinical and vector samples, enabling real-time tracking of viral diversity, evolution, and spread. Experiences from Ebola, Zika, and COVID-19 have demonstrated the power of portable sequencing, now applied to CHIKV monitoring. Advances in tools such as Guppy, Dorado, Minimap2, and Medaka enhance read quality, consensus accuracy, and downstream analyses. Despite challenges in basecalling and error correction, robust quality control pipelines ensure reliable results. Ongoing improvements in chemistry, flow cell design, and machine learning will further enhance fidelity and throughput, establishing nanopore sequencing as a cornerstone of CHIKV genomic surveillance and epidemic preparedness.

Chikungunya virus↗

Micropurification of two human cerebrospinal fluid proteins by high performance electrophoresis chromatography.

Using C8 reversed-phase HPLC in conjunction with sodium dodecyl sulfate-polyacrylamide gel electrophoresis, we have fractionated proteins contained in human CSFs obtained from patients with schizophrenic disorders. When these proteins were electrophoretically blotted onto polyvinylidene difluoride membrane for direct N-terminal amino acid sequencing, several CSF proteins were identified; these included albumin, transferrin, apolipoprotein A-I, beta 2-microglobulin, and prealbumin. We have also identified two structurally related human CSF proteins designated cerebrin 28 (M(r) 28,000) and cerebrin 30 (M(r) 30,000) that have an N-terminal amino acid sequence of NH2-APPAQVSVQPNF and NH2-APEAQVSVQPLFXQ, respectively. Comparison of these sequences with existing database at Protein Identification Resource (R 32.0), GenBank (R 72.0), SWISS-PROT (R 22.0), and EMBL (R 31.0) indicated that they are unique proteins. These proteins were subsequently purified by high performance electrophoresis chromatography (HPEC) using an Applied Biosystems 230A HPEC system. A specific polyclonal antibody was prepared and an ELISA was established for cerebrin 30. It was noted that HPEC is a powerful tool to purify microgram quantities of proteins from human, rabbit, and rat CSFs. Using such a system, we have been able to micropurify as many as 10 proteins simultaneously in a single experiment because the elution of proteins occurred strictly according to their molecular weights. More importantly, we routinely obtained a recovery of > 90%. The potential use of this technology for micropurification of proteins was discussed.

Amino Acid Sequence↗

iProLINK: an integrated protein resource for literature mining.

The exponential growth of large-scale molecular sequence data and of the PubMed scientific literature has prompted active research in biological literature mining and information extraction to facilitate genome/proteome annotation and improve the quality of biological databases. Motivated by the promise of text mining methodologies, but at the same time, the lack of adequate curated data for training and benchmarking, the Protein Information Resource (PIR) has developed a resource for protein literature mining--iProLINK (integrated Protein Literature INformation and Knowledge). As PIR focuses its effort on the curation of the UniProt protein sequence database, the goal of iProLINK is to provide curated data sources that can be utilized for text mining research in the areas of bibliography mapping, annotation extraction, protein named entity recognition, and protein ontology development. The data sources for bibliography mapping and annotation extraction include mapped citations (PubMed ID to protein entry and feature line mapping) and annotation-tagged literature corpora. The latter includes several hundred abstracts and full-text articles tagged with experimentally validated post-translational modifications (PTMs) annotated in the PIR protein sequence database. The data sources for entity recognition and ontology development include a protein name dictionary, word token dictionaries, protein name-tagged literature corpora along with tagging guidelines, as well as a protein ontology based on PIRSF protein family names. iProLINK is freely accessible at http://pir.georgetown.edu/iprolink, with hypertext links for all downloadable files.

Computational Biology↗

Molecular characterization of the aryl hydrocarbon receptors (AHR1 and AHR2) from red seabream (Pagrus major).

The aryl hydrocarbon receptor (AHR) mediates the toxic effects of planar halogenated aromatic hydrocarbons (PHAHs). Bony fishes exposed to PHAHs exhibit a wide range of developmental defects. However, functional roles of fish AHR are not yet fully understood, compared with those of mammalian AHRs. To investigate the potential sensitivity to PHAHs toxic effects, an AHR cDNA was initially cloned and sequenced from red seabream (Pagrus major), an important fishery resource in Japan. The present study succeeded in identifying two highly divergent red seabream AHR cDNA clones, which shared only 32% identity in full-length amino acid sequence. The phylogenetic analysis revealed that one belonged to AHR1 clade (rsAHR1) and another to AHR2 clade (rsAHR2). The rsAHR1 encoded a 846-residue protein with a predicted molecular mass of 93.2 kDa, and 990 amino acids and 108.9 kDa encoded rsAHR2. In the N-terminal half, both rsAHR genes included bHLH and PAS domains, which participate in ligand binding, AHR/ARNT dimerization and DNA binding. The C-terminal half, which is responsible for transactivation, was poorly conserved between rsAHRs. Quantitative analyses of both rsAHRs mRNAs revealed that their tissue expression profiles were isoform-specific; rsAHR1 mRNA expressed primarily in brain, heart, ovary and spleen, while rsAHR2 mRNA was observed in all tissues examined, indicating distinct roles of each rsAHR. Furthermore, there appeared to be species-differences in the tissue expression profiles of AHR isoforms between red seabream and other fish. These results suggest that there are isoform- and species-specific functions in piscine AHRs.

Amino Acid Sequence↗