Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequencing Resource”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Automatic generation of primary sequence patterns from sets of related protein sequences.

We have developed a computer algorithm that can extract the pattern of conserved primary sequence elements common to all members of a homologous protein family. The method involves clustering the pairwise similarity scores among a set of related sequences to generate a binary dendrogram (tree). The tree is then reduced in a stepwise manner by progressively replacing the node connecting the two most similar termini by one common pattern until only a single common "root" pattern remains. A pattern is generated at a node by (i) performing a local optimal alignment on the sequence/pattern pair connected by the node with the use of an extended dynamic programming algorithm and then (ii) constructing a single common pattern from this alignment with a nested hierarchy of amino acid classes to identify the minimal inclusive amino acid class covering each paired set of elements in the alignment. Gaps within an alignment are created and/or extended using a "pay once" gap penalty rule, and gapped positions are converted into gap characters that function as 0 or 1 amino acid of any type during subsequent alignment. This method has been used to generate a library of covering patterns for homologous families in the National Biomedical Research Foundation/Protein Identification Resource protein sequence data base. We show that a covering pattern can be more diagnostic for sequence family membership than any of the individual sequences used to construct the pattern.

Amino Acid Sequence↗

Conifer defence against insects: microarray gene expression profiling of Sitka spruce (Picea sitchensis) induced by mechanical wounding or feeding by spruce budworms (Choristoneura occidentalis) or white pine weevils (Pissodes strobi) reveals large-scale changes of the host transcriptome.

Conifers are resistant to attack from a large number of potential herbivores or pathogens. Previous molecular and biochemical characterization of selected conifer defence systems support a model of multigenic, constitutive and induced defences that act on invading insects via physical, chemical, biochemical or ecological (multitrophic) mechanisms. However, the genomic foundation of the complex defence and resistance mechanisms of conifers is largely unknown. As part of a genomics strategy to characterize inducible defences and possible resistance mechanisms of conifers against insect herbivory, we developed a cDNA microarray building upon a new spruce (Picea spp.) expressed sequence tag resource. This first-generation spruce cDNA microarray contains 9720 cDNA elements representing c. 5500 unique genes. We used this array to monitor gene expression in Sitka spruce (Picea sitchensis) bark in response to herbivory by white pine weevils (Pissodes strobi, Curculionidae) or wounding, and in young shoot tips in response to western spruce budworm (Choristoneura occidentalis, Lepidopterae) feeding. Weevils are stem-boring insects that feed on phloem, while budworms are foliage feeding larvae that consume needles and young shoot tips. Both insect species and wounding treatment caused substantial changes of the host plant transcriptome detected in each case by differential gene expression of several thousand array elements at 1 or 2 d after the onset of treatment. Overall, there was considerable overlap among differentially expressed gene sets from these three stress treatments. Functional classification of the induced transcripts revealed genes with roles in general plant defence, octadecanoid and ethylene signalling, transport, secondary metabolism, and transcriptional regulation. Several genes involved in primary metabolic processes such as photosynthesis were down-regulated upon insect feeding or wounding, fitting with the concept of dynamic resource allocation in plant defence. Refined expression analysis using gene-specific primers and real-time PCR for selected transcripts was in agreement with microarray results for most genes tested. This study provides the first large-scale survey of insect-induced defence transcripts in a gymnosperm and provides a platform for functional investigation of plant-insect interactions in spruce. Induction of spruce genes of octadecanoid and ethylene signalling, terpenoid biosynthesis, and phenolic secondary metabolism are discussed in more detail.

Animals↗

Complexity: an internet resource for analysis of DNA sequence complexity.

The search for DNA regions with low complexity is one of the pivotal tasks of modern structural analysis of complete genomes. The low complexity may be preconditioned by strong inequality in nucleotide content (biased composition), by tandem or dispersed repeats or by palindrome-hairpin structures, as well as by a combination of all these factors. Several numerical measures of textual complexity, including combinatorial and linguistic ones, together with complexity estimation using a modified Lempel-Ziv algorithm, have been implemented in a software tool called 'Complexity' (http://wwwmgs.bionet.nsc.ru/mgs/programs/low_complexity/). The software enables a user to search for low-complexity regions in long sequences, e.g. complete bacterial genomes or eukaryotic chromosomes. In addition, it estimates the complexity of groups of aligned sequences.

Algorithms↗

A BAC-based physical map of the Drosophila buzzatii genome.

Large-insert genomic libraries facilitate cloning of large genomic regions, allow the construction of clone-based physical maps, and provide useful resources for sequencing entire genomes. Drosophila buzzatii is a representative species of the repleta group in the Drosophila subgenus, which is being widely used as a model in studies of genome evolution, ecological adaptation, and speciation. We constructed a Bacterial Artificial Chromosome (BAC) genomic library of D. buzzatii using the shuttle vector pTARBAC2.1. The library comprises 18,353 clones with an average insert size of 152 kb and an approximately 18x expected representation of the D. buzzatii euchromatic genome. We screened the entire library with six euchromatic gene probes and estimated the actual genome representation to be approximately 23x. In addition, we fingerprinted by restriction digestion and agarose gel electrophoresis a sample of 9555 clones, and assembled them using FingerPrint Contigs (FPC) software and manual editing into 345 contigs (mean of 26 clones per contig) and 670 singletons. Finally, we anchored 181 large contigs (containing 7788 clones) to the D. buzzatii salivary gland polytene chromosomes by in situ hybridization of 427 representative clones. The BAC library and a database with all the information regarding the high coverage BAC-based physical map described in this paper are available to the research community.

Animals↗

Large-scale analysis of the barley transcriptome based on expressed sequence tags.

To provide resources for barley genomics, 110,981 expressed sequence tags (ESTs) were generated from 22 cDNA libraries representing tissues at various developmental stages. This EST collection corresponds to approximately one-third of the 380,000 publicly available barley ESTs. Clustering and assembly resulted in 14,151 tentative consensi (TCs) and 11 073 singletons, altogether representing 25 224 putatively unique sequences. Of these, 17.5% showed no significant similarity to other barley ESTs present in dbEST. More than 41% of all barley genes are supposed to belong to multigene families and approximately 4% of the barley genes undergo alternative splicing. Based on the functional annotation of the set of unique sequences, the functional category 'Energy' was further analysed to reveal tissue- and stage-specific differences in gene expression. Hierarchical clustering of 362 differentially expressed TCs resulted in the identification of seven major clusters. The clusters reflect biochemical pathways predominantly activated in specific tissues and at various developmental stages. During seed germination glycolysis could be identified as the most predominant biochemical pathway. Germination-specific glycolysis is characterized by the coordinated expression of phosphoenolpyruvate carboxylase and phosphoenolpyruvate carboxykinase, whose antagonistic actions possibly regulate the flux of amino acids into protein biosynthesis and gluconeogenesis respectively. The expression of defence-related and antioxidant genes during germination might be controlled by the ethylene-signalling pathway as concluded from the coordinated expression of those genes and the transcription factors (TF) EIN3 and EREBPG. Moreover, because of their predominant expression in germinating seeds, TF of the AP2 and MYB type are presumably major regulators of germination.

Expressed Sequence Tags↗

Long-term outcome following case management after coronary artery bypass surgery.

Patient outcome following coronary artery bypass grafting (CABG) has come under increasing governmental, social, and economic scrutiny. To insure quality patient outcome after CABG, many new policies and programs have been instituted. One of these, case management, was developed as a tool for identification and quantification of patient clinical sequences and resource utilization. This present study examines the influence of case management on length of stay and patient outcome following CABG. One hundred forty randomized, retrospectively analyzed CABG patients from 1990, prior to case management, were compared against 140 age-and case-matched randomly controlled CABG patients from 1994 after case management was in place. Patients' demographics were similar. The outcome data showed that intensive care unit (ICU) use and total length of stay were significantly decreased. Furthermore, resource utilization as monitored by chest X-ray, electrocardiography, and laboratory testing were decreased as well. Finally, mortality was decreased despite an increase in risk-adjusted acuity of the patients. There appeared to be no effect of gender or age on the benefit derived from case management. These data demonstrate that the influence of case management is beneficial for resource utilization and patient outcome following CABG and that these types of patient care policy advancements should be encouraged.

Aged↗

Large-scale and automated DNA sequence determination.

DNA sequence analysis is a multistage process that includes the preparation of DNA, its fragmentation and base analysis, and the interpretation of the resulting sequence information. New technological advances have led to the automation of certain steps in this process and have raised the possibility of large-scale DNA sequencing efforts in the near future [for example, 1 million base pairs (Mb) per year]. New sequencing methodologies, fully automated instrumentation, and improvements in sequencing-related computational resources may render genome-size sequencing projects (100 Mb or larger) feasible during the next 5 to 10 years.

Animals↗

[Quality plan of a bone tissue bank].

In 1994, the French bioethics laws changed the regulations concerning donated human tissues with safety precautions and therapeutic use. The decree of 29(th) December 1998 describes the practical obligations required of bone tissue banks for approval by the French Health Ministry. Before the new regulations, a quality system was implemented at the Cochin hospital bone tissue bank. It integrates quality control for each step of the general process. We describe the quality plan of the Cochin Hospital bone tissue bank and show that the specific quality practices, resources and sequence of activities improve distribution of human bone allografts while maintaining their traceability.

Bioethics↗

The PIR-International Protein Sequence Database.

PIR-International is an association of macromolecular sequence data collection centers dedicated to fostering international cooperation as an essential element in the development of scientific databases. A major objective of PIR-International is to continue the development of the Protein Sequence Database as an essential public resource for protein sequence information. This paper briefly describes the architecture of the Protein Sequence Database and how it and associated data sets are distributed and can be accessed electronically.

Amino Acid Sequence↗

[Quality assurance in food production in Europe according to ISO 90000 and HACCP (Hazard Analysis and Critical Control Points)].

HACCP is intended to make food protection programs evolve from a mainly retrospective quality control toward a preventative quality assurance approach and to provide an increased confidence in food safety. In parallel, the continuous evolution of quality concepts in the food industry resulted in the development of quality systems and quality assurance techniques with regard to the EN 29,000 (ISO 9000) series of standards. Specific to the food industry, HACCP can be seen as a very effective method to prepare specific Safety Assurance Plans (cf. the quality assurance plan concept) within a quality systems approach (Jouve, 1993). By reference to ISO 8402, a Quality Assurance Plan (QAP) basically sets out "the specific quality practices, resources and sequence of activities relevant to a particular product, service, contract or project". Quality assurance plans are more particularly useful for projects relating to new products or processes and/or comprising inter-related tasks whose interaction may be complex. In addition, in contractual or regulatory situations, such plans can be used to demonstrate the supplier's capability to meet identified objectives, specification or standards. In the food industry, the management of safety is a critical and complex issue which fits very well in the scope of application of a specific QAP; it is also where the use of HACCP is otherwise recommended by priority.

Food Contamination↗

Sequence-tagged connectors: a sequence approach to mapping and scanning the human genome.

The sequence-tagged connector (STC) strategy proposes to generate sequence tags densely scattered (every 3.3 kilobases) across the human genome by arraying 450,000 bacterial artificial chromosomes (BACs) with randomly cleaved inserts, sequencing both ends of each, and preparing a restriction enzyme fingerprint of each. The STC resource, containing end sequences, fingerprints, and arrayed BACs, creates a map where the interrelationships of the individual BAC clones are resolved through their STCs as overlapping BAC clones are sequenced. Once a seed or initiation BAC clone is sequenced, the minimum overlapping 5' and 3' BAC clones can be identified computationally and sequenced. By reiterating this "sequence-then-map by computer analysis against the STC database" strategy, a minimum tiling path of clones can be sequenced at a rate that is primarily limited by the sequencing throughput of individual genome centers. As of February 1999, we had deposited, together with The Institute for Genomic Research (TIGR), into GenBank 314,000 STCs ( approximately 135 megabases), or 4.5% of human genomic DNA. This genome survey reveals numerous genes, genome-wide repeats, simple sequence repeats (potential genetic markers), and CpG islands (potential gene initiation sites). It also illustrates the power of the STC strategy for creating minimum tiling paths of BAC clones for large-scale genomic sequencing. Because the STC resource permits the easy integration of genetic, physical, gene, and sequence maps for chromosomes, it will be a powerful tool for the initial analysis of the human genome and other complex genomes.

Chromosome Mapping↗

From mapping to sequencing, post-sequencing and beyond.

The Rice Genome Research Program (RGP) in Japan has been collaborating with the international community in elucidating a complete high-quality sequence of the rice genome. As the pioneer in large-scale analysis of the rice genome, the RGP has successfully established the fundamental tools for genome research such as a genetic map, a yeast artificial chromosome (YAC)-based physical map, a transcript map and a phage P1 artificial chromosome (PAC)/bacterial artificial chromosome (BAC) sequence-ready physical map, which serve as common resources for genome sequencing. Among the 12 rice chromosomes, the RGP is in charge of sequencing six chromosomes covering 52% of the 390 Mb total length of the genome. The contribution of the RGP to the realization of decoding the rice genome sequence with high accuracy and deciphering the genetic information in the genome will have a great impact in understanding the biology of the rice plant that provides a major food source for almost half of the world's population. A high-quality draft sequence (phase 2) was completed in December 2002. Since then, much of the finished quality sequence (phase 3) has become available in public databases. With the completion of sequencing in December 2004, it is expected that the genome sequence would facilitate innovative research in functional and applied genomics. A map-based genome sequence is indispensable for further improvement of current rice varieties and for development of novel varieties carrying agronomically important traits such as high yield potential and tolerance to both biotic and abiotic stresses. In addition to genome sequencing, various related projects have been initiated to generate valuable resources, which could serve as indispensable tools in clarifying the structure and function of the rice genome. These resources have been made available to the scientific community through the Rice Genome Resource Center (RGRC) of the National Institute of Agrobiological Sciences (NIAS) to enable rapid progress in research that will lead to thorough understanding of the rice plant. As the next trend in rice genome research will focus on determining the function of about 40,000-50,000 genes predicted in the genome as well as applying various genomics tools in rice breeding, an unlimited access to rice DNA and seed stocks will provide a broad community of scientists with the necessary materials for formulating new concepts, developing innovative research and making new scientific discoveries in rice genomics.

Centromere↗

Modeling sequencing errors by combining Hidden Markov models.

Among the largest resources for biological sequence data is the large amount of expressed sequence tags (ESTs) available in public and proprietary databases. ESTs provide information on transcripts but for technical reasons they often contain sequencing errors. Therefore, when analyzing EST sequences computationally, such errors must be taken into account. Earlier attempts to model error prone coding regions have shown good performance in detecting and predicting these while correcting sequencing errors using codon usage frequencies. In the research presented here, we improve the detection of translation start and stop sites by integrating a more complex mRNA model with codon usage bias based error correction into one hidden Markov model (HMM), thus generalizing this error correction approach to more complex HMMs. We show that our method maintains the performance in detecting coding sequences.

Algorithms↗

SynBrowse: a synteny browser for comparative sequence analysis.

MOTIVATION: The recent efforts of various sequence projects to sequence deeply into various phylogenies provide great resources for comparative sequence analysis. A generic and portable tool is essential for scientists to visualize and analyze sequence comparisons. RESULTS: We have developed SynBrowse, a synteny browser for visualizing and analyzing genome alignments both within and between species. It is intended to help scientists study macrosynteny, microsynteny and homologous genes between sequences. It can also aid with the identification of uncharacterized genes, putative regulatory elements and novel structural features of a species. SynBrowse is a GBrowse (the Generic Genome Browser) family software tool that runs on top of the open source BioPerl modules. It consists of two components: a web-based front end and a set of relational database back ends. Each database stores pre-computed alignments from a focus sequence to reference sequences in addition to the genome annotations of the focus sequence. The user interface lets end users select a key comparative alignment type and search for syntenic blocks between two sequences and zoom in to view the relationships among the corresponding genome annotations in detail. SynBrowse is portable with simple installation, flexible configuration, convenient data input and easy integration with other components of a model organism system. AVAILABILITY: The software is available at http://www.gmod.org CONTACT: vbrendel@iastate.edu

Algorithms↗

The first filamentous fungal genome sequences: Aspergillus leads the way for essential everyday resources or dusty museum specimens?

The published Aspergillus genome sequences (A. nidulans, A. fumigatus, A. oryzae) and further sequence data from A. clavatus, Neosartorya fischeri, A. flavus, A. niger, A. parasiticus and A. terreus are the first from a group of related filamentous fungi. They indicate the gains possible from genomic approaches, but also problems that arise after the sequences are finished. Benefits include a greater understanding of genome structure and evolution, insights into gene regulation, predictions of new factors that may be relevant to pathogenicity and the discovery of novel enzymes with biotechnological value. Areas where further developments are needed include gene and structure-function predictions, methods for comparative genome analysis and the interfaces for access to genome information. In addition, strategies for continued maintenance and updating need to be developed at the start of the post-genomic era to increase the value of genome sequences into the future.

Aspergillus↗

Public databases: retrieving and manipulating sequences for beginners.

This chapter outlines the basic requirements for finding and exploring sequences of interest in public databases, such as GenBank. As such, it is not aimed at experienced sequencers, for whom this will be "second nature," but at the many clinical bacteriologists who rarely have need of DNA sequences in their usual work, and who would like to develop their interest in what can appear to be a daunting area. The topics discussed include finding and retrieving sequences from GenBank, identifying homologous sequences using BLAST searches, resources for accessing microbial genomes, and the Protein Data Bank. Finally, recommendations are made for useful software (freeware) and online sequence manipulation resources.

Amino Acid Sequence↗

Human cortical networks for new and familiar sequences of saccades.

Visual exploration is organized in sequences of saccadic eye movements that depend on both perceptual and cognitive context. Using functional magnetic resonance imaging, we studied the neural basis of sequential oculomotor behavior and its dependence on different types of memory by analyzing cerebral activity during performance of newly learned and familiar sequences of eye movements. Compared to a resting condition, both types of sequences activated a common fronto-parietal network, including frontal and supplementary eye fields, and several parietal areas. Within this network, newly learned sequences induced stronger activation than familiar sequences, probably reflecting higher attentional demands. In addition, specific regions were recruited for the performance of new sequences, including pre-supplementary eye fields, the precuneus and the caudate nucleus. This indicates that in addition to attentional modulation, novelty of saccadic sequences requires specific cortical resources, probably related to effortful sequence preparation and coordination as well as to spatial working memory. For familiar sequences, recalled from long-term memory, we observed specific right medial temporo-occipital activation in the vicinity of the boundary between the parahippocampal and lingual gyri, as well as an activation site in the parieto-occipital fissure. We conclude that neuronal resources recruited by the gaze system can change with the familiarity of the scanpath to be executed. This study is important to better understand how the brain implements memorized scanpaths for visual exploration and orienting.

Adult↗

A sea urchin genome project: sequence scan, virtual map, and additional resources.

Results of a first-stage Sea Urchin Genome Project are summarized here. The species chosen was Strongylocentrotus purpuratus, a research model of major importance in developmental and molecular biology. A virtual map of the genome was constructed by sequencing the ends of 76,020 bacterial artificial chromosome (BAC) recombinants (average length, 125 kb). The BAC-end sequence tag connectors (STCs) occur an average of 10 kb apart, and, together with restriction digest patterns recorded for the same BAC clones, they provide immediate access to contigs of several hundred kilobases surrounding any gene of interest. The STCs survey >5% of the genome and provide the estimate that this genome contains approximately 27,350 protein-coding genes. The frequency distribution and canonical sequences of all middle and highly repetitive sequence families in the genome were obtained from the STCs as well. The 500-kb Hox gene complex of this species is being sequenced in its entirety. In addition, arrayed cDNA libraries of >10(5) clones each were constructed from every major stage of embryogenesis, several individual cell types, and adult tissues and are available to the community. The accumulated STC data and an expanding expressed sequence tag database (at present including >12, 000 sequences) have been reported to GenBank and are accessible on public web sites.

Aging↗