Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequencing Resource”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Development and mapping of SSR markers for maize.

Microsatellite or simple sequence repeat (SSR) markers have wide applicability for genetic analysis in crop plant improvement strategies. The objectives of this project were to isolate, characterize, and map a comprehensive set of SSR markers for maize (Zea mays L.). We developed 1051 novel SSR markers for maize from microsatellite-enriched libraries and by identification of microsatellite-containing sequences in public and private databases. Three mapping populations were used to derive map positions for 978 of these markers. The main mapping population was the intermated B73 x Mo17 (IBM) population. In mapping this intermated recombinant inbred line population, we have contributed to development of a new high-resolution map resource for maize. The primer sequences, original sequence sources, data on polymorphisms across 11 inbred lines, and map positions have been integrated with information on other public SSR markers and released through MaizeDB at URL:www.agron.missouri.edu. The maize research community now has the most detailed and comprehensive SSR marker set of any plant species.

Chromosome Mapping↗

A statistical model for HIV-1 sequence classification using the subtype analyser (STAR).

MOTIVATION: HIV-1 antiretroviral drug resistance testing produces large amounts of HIV-1 protease and reverse transcriptase sequences. These provide an excellent resource to study the incidence, spread and clinical significance of HIV-1 subtypes. We have produced a program, Subtype Analyser (STAR) that rapidly and accurately subtypes HIV-1. Here we have determined a robust and statistically validated model for subtype assignment. RESULTS: We have significantly extended our HIV-1 subtyping tool (STAR), such that each query sequence when evaluated against subtype profile alignments, returns a discriminating score based on the ratio of subtype positive to negative amino acid positions. These scores were transformed into a Z-score distribution and evaluated. Of the 141 sequences used to define the subtype alignments, 98% were correctly reclassified. Inclusion of additional recombination detection within STAR increased the detection of known recombinant sequences to 95%. AVAILABILITY: STAR is available as compiled (Linux Fedora 3) or source code from http://pgv19.virol.ucl.ac.uk/download/star_linux.tar CONTACT: p.kellam@ucl.ac.uk SUPPLEMENTARY INFORMATION: http://pgv19.virol.ucl.ac.uk/download/star_supplement

Algorithms↗

Sequencing and analysis of common bean ESTs. Building a foundation for functional genomics.

Although common bean (Phaseolus vulgaris) is the most important grain legume in the developing world for human consumption, few genomic resources exist for this species. The objectives of this research were to develop expressed sequence tag (EST) resources for common bean and assess nodule gene expression through high-density macroarrays. We sequenced a total of 21,026 ESTs derived from 5 different cDNA libraries, including nitrogen-fixing root nodules, phosphorus-deficient roots, developing pods, and leaves of the Mesoamerican genotype, Negro Jamapa 81. The fifth source of ESTs was a leaf cDNA library derived from the Andean genotype, G19833. Of the total high-quality sequences, 5,703 ESTs were classified as singletons, while 10,078 were assembled into 2,226 contigs producing a nonredundant set of 7,969 different transcripts. Sequences were grouped according to 4 main categories, metabolism (34%), cell cycle and plant development (11%), interaction with the environment (19%), and unknown function (36%), and further subdivided into 15 subcategories. Comparisons to other legume EST projects suggest that an entirely different repertoire of genes is expressed in common bean nodules. Phaseolus-specific contigs, gene families, and single nucleotide polymorphisms were also identified from the EST collection. Functional aspects of individual bean organs were reflected by the 20 contigs from each library composed of the most redundant ESTs. The abundance of transcripts corresponding to selected contigs was evaluated by RNA blots to determine whether gene expression determined by laboratory methods correlated with in silico expression. Evaluation of root nodule gene expression by macroarrays and RNA blots showed that genes related to nitrogen and carbon metabolism are integrated for ureide production. Resources developed in this project provide genetic and genomic tools for an international consortium devoted to bean improvement.

Carbon↗

Gene organization in the bleomycin-resistance region of the producer organism Streptomyces verticillus.

A nucleotide sequence of 7 kb is reported, encompassing two bleomycin-resistance (BmR-encoding) genes and five other open reading frames (ORFs) from the Bm-producing organism Streptomyces verticillus ATCC 15003. The deduced ORFs, in sequence order, encode for (i) a protein homologous to an amino-acid dioxygenase; (ii) BlmA, the BmR-binding protein described by Sugiyama et al. [Gene 151 (1994) 11-16]; (iii) a product containing three copies of a sequence homologous to the ankyrin repeat; (iv) a product lacking homology to any of the sequences in the Protein Identification Resource database (PIR), release 37; (v) BlmB, the BmR acetyltransferase described by Sugiyama et al. (1994); (vi) an unidentified protein which augmented resistance determined by ORF2 (BlmA); (vii) a member of the ATP-binding cassette (ABC) family of transport protein. Predicted translational frameshifts in the -1 frame occur at the junctions between ORF3 and ORF4, ORF4 and ORF5, and ORF6 and ORF7. Sequences homologous to ORF2 and ORF3 were identified in the genome of the producer organism for the related antibiotic phleomycin.

Acetyltransferases↗

Putative rat sperm lipid-binding protein: isolation and partial characterization.

Previous work has identified a prominent 22-24-kD protein that is present in rat male reproductive tissues, including epididymis and testis (Brooks, 1985; Jones and Brown, 1987; Moore et al., 1987). Using a monoclonal antibody (designated mAb-B109) against this 24-kD antigen (referred to as B109), we have isolated the protein using a combination of chromatofocusing and electroelution from SDS-PAGE gels, and reverse phase HPLC. B109 (pI = 4.8) is amino-terminal blocked. To obtain internal amino acid sequences, the isolated protein was cleaved either with cyanogen bromide in 70% formic acid or with TLCK-treated chymotrypsin. With cyanogen bromide treatment, two peptides, 17.8 kD and 11.9 kD, were isolated and partial amino acid sequences obtained. Chymotryptic peptides were isolated by reverse-phase HPLC and two were chosen for sequence analysis. A computer search for sequence homology through the protein identification resource (PIR) matched B109 to a basic 21-kD cytosolic protein (pI = 7.4) found in bovine brain (> 80% homology). When peptide sequence differences obtained in the present study were substituted into the 21-kD cytosolic protein sequence obtained from the PIR using Intelligenetics software, the calculated pI dropped from 7.4 to 5.8, suggesting that pI differences between the bovine and rat molecules are the result of amino acid substitutions in the testis protein and not tissue-specific posttranslational processing. It has been postulated that the 21-kD bovine brain protein is associated with phospholipid transport, although the function of B109 is unknown.

Amino Acid Sequence↗

A genome-wide screen for normally methylated human CpG islands that can identify novel imprinted genes.

DNA methylation is a covalent modification of the nucleotide cytosine that is stably inherited at the dinucleotide CpG by somatic cells, and 70% of CpG dinucleotides in the genome are methylated. The exception to this pattern of methylation are CpG islands, CpG-rich sequences that are protected from methylation, and generally are thought to be methylated only on the inactive X-chromosome and in tumors, as well as differentially methylated regions (DMRs) in the vicinity of imprinted genes. To identify chromosomal regions that might harbor imprinted genes, we devised a strategy for isolating a library of normally methylated CpG islands. Most of the methylated CpG islands represented high copy number dispersed repeats. However, 62 unique clones in the library were characterized, all of which were methylated and GC-rich, with a GC content >50%. Of these, 43 clones also showed a CpG(obs)/CpG(exp) >0.6, of which 30 were studied in detail. These unique methylated CpG islands mapped to 23 chromosomal regions, and 12 were differentially methylated regions in uniparental tissues of germline origin, i.e., hydatidiform moles (paternal origin) and complete ovarian teratomas (maternal origin), even though many apparently were methylated in somatic tissues. We term these sequences gDMRs, for germline differentially methylated regions. At least two gDMRs mapped near imprinted genes, HYMA1 and a novel homolog of Elongin A and Elongin A2, which we term Elongin A3. Surprisingly, 18 of the methylated CpG islands were methylated in germline tissues of both parental origins, representing a previously uncharacterized class of normally methylated CpG islands in the genome, and which we term similarly methylated regions (SMRs). These SMRs, in contrast to the gDMRs, were significantly associated with telomeric band locations (P =.0008), suggesting a potential role for SMRs in chromosome organization. At least 10 of the methylated CpG islands were on average 85% conserved between mouse and human. These sequences will provide a valuable resource in the search for novel imprinted genes, for defining the molecular substrates of the normal methylome, and for identifying novel targets for mammalian chromatin formation.

Amino Acid Sequence↗

Genetics of the immune response: identifying immune variation within the MHC and throughout the genome.

With the advent of modern genomic sequencing technology the ability to obtain new sequence data and to acquire allelic polymorphism data from a broad range of samples has become routine. In this regard, our investigations have started with the most polymorphic of genetic regions fundamental to the immune response in the major histocompatibility complex (MHC). Starting with the completed human MHC genomic sequence, we have developed a resource of methods and information that provide ready access to a large portion of human and nonhuman primate MHCs. This resource consists of a set of primer pairs or amplicons that can be used to isolate about 15% of the 4.0 Mb MHC. Essentially similar studies are now being carried out on a set of immune response loci to broaden the usefulness of the data and tools developed. A panel of 100 genes involved in the immune response have been targeted for single nucleotide polymorphism (SNP) discovery efforts that will analyze 120 Mb of sequence data for the presence of immune-related SNPs. The SNP data provided from the MHC and from the immune response panel has been adapted for use in studies of evolution, MHC disease associations, and clinical transplantation.

Animals↗

Mouse BAC ends quality assessment and sequence analyses.

A large-scale BAC end-sequencing project at The Institute for Genomic Research (TIGR) has generated one of the most extensive sets of sequence markers for the mouse genome to date. With a sequencing success rate of >80%, an average read length of 485 bp, and ABI3700 capillary sequencers, we have generated 449,234 nonredundant mouse BAC end sequences (mBESs) with 218 Mb total from 257,318 clones from libraries RPCI-23 and RPCI-24, representing 15x clone coverage, 7% sequence coverage, and a marker every 7 kb across the genome. A total of 191,916 BACs have sequences from both ends providing 12x genome coverage. The average Q20 length is 406 bp and 84% of the bases have phred quality scores > or = 20. RPCI-24 mBESs have more Q20 bases and longer reads on average than RPCI-23 sequences. ABI3700 sequencers and the sample tracking system ensure that > 95% of mBESs are associated with the right clone identifiers. We have found that a significant fraction of mBESs contains L1 repeats and approximately 48% of the clones have both ends with > or = 100 bp contiguous unique Q20 bases. About 3% mBESs match ESTs and > 70% of matches were conserved between the mouse and the human or the rat. Approximately 0.1% mBESs contain STSs. About 0.2% mBESs match human finished sequences and > 70% of these sequences have EST hits. The analyses indicate that our high-quality mouse BAC end sequences will be a valuable resource to the community.

Animals↗

Allocation of attention to programming of movement sequences in Parkinson's disease.

The allocation of attention to the programming and execution of movement sequences was examined in Parkinson's disease (PD). The time taken to initiate and execute sequences of one, three, and five button taps was examined, while also varying the hand used (left or right) and the attentional resources that could be allocated to sequencing (using single- versus dual-task conditions). These results showed that performance anomalies in PD were most apparent with the preferred right hand under single-rather than dual-task conditions. Subjects suffering from PD may tend to divert attention from the right hand under single-task conditions, and perhaps with short sequences, as well as being less likely to prepare sequences of more than three movements in advance with that hand. These effects were unlikely to reflect asymmetric pathology. If the right hand of such subjects has in some respects now come to behave more like a "clumsy" left hand, this may reflect a deliberate strategic choice in an attempt to cope with a movement impairment.

Aged↗

Fungal BLAST and Model Organism BLASTP Best Hits: new comparison resources at the Saccharomyces Genome Database (SGD).

The Saccharomyces Genome Database (SGD; http://www.yeastgenome.org/) is a scientific database of gene, protein and genomic information for the yeast Saccharomyces cerevisiae. SGD has recently developed two new resources that facilitate nucleotide and protein sequence comparisons between S.cerevisiae and other organisms. The Fungal BLAST tool provides directed searches against all fungal nucleotide and protein sequences available from GenBank, divided into categories according to organism, status of completeness and annotation, and source. The Model Organism BLASTP Best Hits resource displays, for each S.cerevisiae protein, the single most similar protein from several model organisms and presents links to the database pages of those proteins, facilitating access to curated information about potential orthologs of yeast proteins.

Databases, Genetic↗

A survey of nucleic acid services in core laboratories.

Core facility services related to DNA synthesis and sequencing were surveyed by the Association of Biomolecular Resource Facilities. Responses from 85 facilities offering DNA synthesis and 37 facilities offering DNA sequencing were obtained. Data on instrumentation, volume, number of users, cost, methodology and a number of other criteria were obtained. The volume of work performed by these centralized core facilities was quite substantial (combined synthesis output of 4 million bases per year and a combined sequencing output of 35 million bases per year). The large number of users supported by these facilities and the high sample throughput make these core resource facilities good indicators of technological trends.

Costs and Cost Analysis↗

ONT-only genome assembly of a Korean male individual using a semen sample.

BACKGROUND: Long-read sequencing has enabled the generation of high-quality human genome assemblies, but many previous assemblies were based on blood-derived DNA and often relied on limited data types from a single sequencing strategy. OBJECTIVE: This study aimed to generate high-quality phased genome assemblies of a Korean individual using multiple independent long-read datasets produced from a single sequencing platform and to evaluate their utility for chromosome-scale assembly and variant detection. METHODS: Genomic DNA was extracted from a semen sample of a Korean male. Long-read, ultra-long-read, and chromatin conformation capture sequencing data were generated using Oxford Nanopore Technologies. These datasets were integrated to construct phased genome assemblies, followed by correction of noticeable phasing errors and assessment of assembly continuity, chromosomal representation, telomeric repeat recovery, and variant detection performance. RESULTS: The final phased assemblies spanned approximately 2.9 Gb and represented 23 pairs of chromosomes with an NG50 of 150 Mb. Telomeric repeats were detected at 36 and 37 of the 48 chromosomal ends in the two assemblies, indicating high end-to-end completeness. In addition, we successfully identified structural variants, including small variants. These results demonstrate that combining multiple Oxford Nanopore data types can produce highly continuous and informative phased human genome assemblies. CONCLUSIONS: We generated high-quality phased genome assemblies of a Korean individual using Oxford Nanopore long-read sequencing data derived from semen DNA. This publicly available genome resource will support broader applications of long-read sequencing in human genomics and variant analysis.

Humans↗

A bacterial artificial chromosome library for sequencing the complete human genome.

A 30-fold redundant human bacterial artificial chromosome (BAC) library with a large average insert size (178 kb) has been constructed to provide the intermediate substrate for the international genome sequencing effort. The DNA was obtained from a single anonymous volunteer, whose identity was protected through a double-blind donor selection protocol. DNA fragments were generated by partial digestion with EcoRI (library segments 1--4: 24-fold) and MboI (segment 5: sixfold) and cloned into the pBACe3.6 and pTARBAC1 vectors, respectively. The quality of the library was assessed by extensive analysis of 169 clones for rearrangements and artifacts. Eighteen BACs (11%) revealed minor insert rearrangements, and none was chimeric. This BAC library, designated as "RPCI-11," has been used widely as the central resource for insert-end sequencing, clone fingerprinting, high-throughput sequence analysis and as a source of mapped clones for diagnostic and functional studies.

Chromosomes, Artificial, Bacterial↗

DirectASRM: uncovering allele-specific post-transcriptional RNA modifications through direct RNA sequencing.

SUMMARY: We developed DirectASRM, a comprehensive database for the systematic identification, integration, and annotation of allele-specific RNA modifications (ASRMs) from direct RNA sequencing data. DirectASRM enables single-base, transcript-level detection of ASRMs across multiple RNA modification types, diverse organisms and condition-specific contexts. The database further evaluates the confidence of each ASRM-SNP pair association within isoform context by jointly considering statistical evidence of allelic modification imbalance and independent support from external next-generation sequencing (NGS) - based RNA modification resources. DirectASRM also provides extensive functional annotations for ASRMs and their associated variants, including intra-sample transcript-level allele-specific expression (ASE) and allele-specific splicing, as well as additional post-transcriptional regulatory features such as miRNA binding, circRNA, RNA-protein interactions, and disease relevance. Overall, DirectASRM serves as a comprehensive resource that supports systematic investigation of the potential functional impact of genetic variants in epitranscriptomic regulation. AVAILABILITY AND IMPLEMENTATION: DirectASRM database is freely accessible at http://modinfor.com/DirectASRM/. DirectASRM pipeline is available at GitHub (https://github.com/jiayin1101/DirectASRM_pipeline) and Zenodo (DOI: https://doi.org/10.5281/zenodo.19876077).

Alleles↗

EST-PAC a web package for EST annotation and protein sequence prediction.

With the decreasing cost of DNA sequencing technology and the vast diversity of biological resources, researchers increasingly face the basic challenge of annotating a larger number of expressed sequences tags (EST) from a variety of species. This typically consists of a series of repetitive tasks, which should be automated and easy to use. The results of these annotation tasks need to be stored and organized in a consistent way. All these operations should be self-installing, platform independent, easy to customize and amenable to using distributed bioinformatics resources available on the Internet. In order to address these issues, we present EST-PAC a web oriented multi-platform software package for expressed sequences tag (EST) annotation. EST-PAC provides a solution for the administration of EST and protein sequence annotations accessible through a web interface. Three aspects of EST annotation are automated: 1) searching local or remote biological databases for sequence similarities using Blast services, 2) predicting protein coding sequence from EST data and, 3) annotating predicted protein sequences with functional domain predictions. In practice, EST-PAC integrates the BLASTALL suite, EST-Scan2 and HMMER in a relational database system accessible through a simple web interface. EST-PAC also takes advantage of the relational database to allow consistent storage, powerful queries of results and, management of the annotation process. The system allows users to customize annotation strategies and provides an open-source data-management environment for research and education in bioinformatics.

Journal Article↗

BAliBASE (Benchmark Alignment dataBASE): enhancements for repeats, transmembrane sequences and circular permutations.

BAliBASE is specifically designed to serve as an evaluation resource to address all the problems encountered when aligning complete sequences. The database contains high quality, manually constructed multiple sequence alignments together with detailed annotations. The alignments are all based on three-dimensional structural superpositions, with the exception of the transmembrane sequences. The first release provided sets of reference alignments dealing with the problems of high variability, unequal repartition and large N/C-terminal extensions and internal insertions. Here we describe version 2.0 of the database, which incorporates three new reference sets of alignments containing structural repeats, trans-membrane sequences and circular permutations to evaluate the accuracy of detection/prediction and alignment of these complex sequences. BAliBASE can be viewed at the web site http://www-igbmc.u-strasbg. fr/BioInfo/BAliBASE2/index.html or can be downloaded from ftp://ftp-igbmc.u-strasbg.fr/pub/BAliBASE2 /.

Algorithms↗

DOCKGROUND resource for studying protein-protein interfaces.

MOTIVATION: Public resources for studying protein interfaces are necessary for better understanding of molecular recognition and developing intermolecular potentials, search procedures and scoring functions for the prediction of protein complexes. RESULTS: The first release of the DOCKGROUND resource implements a comprehensive database of co-crystallized (bound-bound) protein-protein complexes, providing foundation for the upcoming expansion to unbound (experimental and simulated) protein-protein complexes, modeled protein-protein complexes and systematic sets of docking decoys. The bound-bound part of DOCKGROUND is a relational database of annotated structures based on the Biological Unit file (Biounit) provided by the RCSB as a separated file containing probable biological molecule. DOCKGROUND is automatically updated to reflect the growth of PDB. It contains 67,220 pairwise complexes that rely on 14,913 Biounit entries from 34,778 PDB entries (January 30, 2006). The database includes a dynamic generation of non-redundant datasets of pairwise complexes based either on the structural similarity (SCOP classification) or on user-defined sequence identity. The growing DOCKGROUND resource is designed to become a comprehensive public environment for developing and validating new methodologies for modeling of protein interactions. AVAILABILITY: DOCKGROUND is available at http://dockground.bioinformatics.ku.edu. The current first release implements the bound-bound part.

Binding Sites↗

An integrated approach for comparative mapping in rice and barley with special reference to the Rph16 resistance locus.

The accumulated sequence information of the almost completed rice genome and the transcriptome of other cereals provide an excellent starting point for comparative genome analysis. We performed targeted synteny-based marker saturation for the Rph16 leaf rust resistance locus in barley by extensively exploiting these newly available resources. Out of a collection of over 320,000 public barley ESTs 309 non-redundant candidate syntenic clones have been identified for this region in a two-step in silico selection procedure. For mapping, 54 barley cDNA-clones were selected due to the even distribution of their homologs on a putatively collinear 3-Mb rice BAC contig. Out of these, 97% (30) of the polymorphic markers could be genetically assigned in collinearity to the target region in barley and a set of 11 markers was integrated into an rph16 high-resolution map. Although, the collinear target region of rice does not contain an obvious candidate gene for rph16 the results demonstrate the potential of the presented procedure to efficiently utilize EST resources for synteny-based marker saturation. The systematic genome-wide exploitation of the increasing sequence data resources will strongly improve our current view of genome conservation and likely facilitate a synteny-based isolation of genes conserved across cereal species.

Chromosome Mapping↗