Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequencing Resource”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Protein classification artificial neural system.

A neural network classification method is developed as an alternative approach to the large database search/organization problem. The system, termed Protein Classification Artificial Neural System (ProCANS), has been implemented on a Cray supercomputer for rapid superfamily classification of unknown proteins based on the information content of the neural interconnections. The system employs an n-gram hashing function that is similar to the k-tuple method for sequence encoding. A collection of modular back-propagation networks is used to store the large amount of sequence patterns. The system has been trained and tested with the first 2,148 of the 8,309 entries of the annotated Protein Identification Resource protein sequence database (release 29). The entries included the electron transfer proteins and the six enzyme groups (oxidoreductases, transferases, hydrolases, lyases, isomerases, and ligases), with a total of 620 superfamilies. After a total training time of seven Cray central processing unit (CPU) hours, the system has reached a predictive accuracy of 90%. The classification is fast (i.e., 0.1 Cray CPU second per sequence), as it only involves a forward-feeding through the networks. The classification time on a full-scale system embedded with all known superfamilies is estimated to be within 1 CPU second. Although the training time will grow linearly with the number of entries, the classification time is expected to remain low even if there is a 10-100-fold increase of sequence entries. The neural database, which consists of a set of weight matrices of the networks, together with the ProCANS software, can be ported to other computers and made available to the genome community. The rapid and accurate superfamily classification would be valuable to the organization of protein sequence databases and to the gene recognition in large sequencing projects.

Computers, Mainframe↗

[Comparative analysis of sequences of the 5S rDNA NTS in wild close relatives of barley from Tibet of China].

Tibet is a center of distribution and differentiation of genus Hordeum in China,and has a great deal of species resource. The sequences of the nontranscribed intergenic spacer (NTS) region of 5S nuclear ribosomal DNA was studied in 18 varieties of the close relative of wild barley (Hordeum spontaneum) from different geographical regions in Tibet and Turkmenistan. These sequences were determined by sequencing the clones of PCR products. Alignment of sequences revealed that the 5S rDNA NTS contained two comparatively conservative regions, A and B, and a variable TAG repeat (V). The number of TAG repeats varies from 4-17, also including several transitions and conversions (TAG-->TCG, TAG-->TAC). The total size of the conservative regions was from 168 -169 bp, and the sequence length variation was only 1 bp. The GC content (%) of the conservative sequence was 43.8% and the homologous of that nearly 98.2% -100%. The number of variable sites was 12 (7.1%). In general, there were more transitions than conversions in the variation sites,and the ratios of transition/conversion were 1.0 - 2.0. NTS polymorphism of 5S rDNA was mainly determined by polymorphism of TAG repeats. The molecular system-tree showed that the NTS should be a useful marker to classify the accessions in Hordeum.

Base Sequence↗

Systems-wide chicken DNA microarrays, gene expression profiling, and discovery of functional genes.

The goal of our current consortium project is to launch a new era--functional genomics of poultry--by providing genomic resources [expressed sequence tags (EST) and DNA microarrays] and by examining global gene expression in target tissues of chickens. DNA microarray analysis has been a fruitful strategy for the identification of functional genes in several model organisms (i.e., human, rodents, fruit fly, etc.). We have constructed and normalized five tissue-specific or multiple-tissue chicken cDNA libraries [liver, fat, breast, and leg muscle/epiphyseal growth plate, pituitary/hypothalamus/pineal, and reproductive tract (oviduct/ovary/testes)] for high-throughput DNA sequencing of EST. DNA sequence clustering was used to build contigs of overlapping sequence and to identify unique, non-redundant EST clones (unigenes), which permitted printing of systems-wide chicken DNA microarrays. One of the most promising genetic resources for gene exploration and functional gene mapping is provided by two sets of experimental lines of broiler-type chickens developed at INRA, France, by divergent selection for extremes in growth traits (fast-growing versus slow-growing; fatness versus leanness at a similar growth rate). We are using DNA microarrays for global gene expression profiling to identify candidate genes and to map growth, metabolic, and regulatory pathways that control important production traits. Candidate genes will be used for functional gene mapping and QTL analysis of F2 progeny from intercrosses made between divergent genetic lines (fat x lean lines; fast-growing x slow-growing lines). Using our first chicken liver microarray, we have already identified several interesting differentially expressed genes in commercial broilers and in divergently selected broiler lines. Many of these candidate genes are involved in the lipogenic pathway and are controlled in part by the thyrotropic axis. Thus, genome-wide transcriptional profiling is a powerful tool used to visualize the cascade of genetic circuits that govern complex biological responses. Global gene expression profiling and QTL scans should enable us to functionally map the genetic pathways that control growth, development, and metabolism of chickens. This emerging technology will have broad applications for poultry breeding programs (i.e., use of molecular markers) and for future production systems (i.e., the health and welfare of birds and the quality of poultry products).

Animal Husbandry↗

[Sequence of the ITS region of nuclear ribosomal DNA(nrDNA) in Xinjiang wild Dianthus and its phylogenetic relationship].

Xinjiang is a center of distribution and differentiation of genus Dianthus in China, and has a great deal of species resources. The sequences of ITS region (including ITS-1, 5.8S rDNA and ITS-2) of nuclear ribosomal DNA from 8 species of genus Dianthus wildly distributed in Xinjiang were determined by direct sequencing of PCR products. The result showed that the size of the ITS of Dianthus is from 617 to 621 bp, and the length variation is only 4 bp. There are very high homogeneous (97.6%-99.8%) sequences between species, and about 80% homogeneous sequences between genus Dianthus and outgroup. The sequences of ITS in genus Dianthus are relatively conservative. In general, there are more conversion than transition in the variation sites among genus Dianthus. The conversion rates are relatively high, and the ratios of conversion/transition are 1.0-3.0. On the basis of phylogenetic analysis of nucleotide sequences the species of Dianthus in China would be divided into three sections. There is a distant relationship between sect. Barbulatum Williams and sect. Dianthus and between sect. Barbulatum Williams and sect. Fimbriatum Williams, and there is a close relationship between sect. Dianthus and sect. Fimbriatum Williams. From the phylogenetic tree of ITS it was found that the origin of sect. Dianthusis is earlier than that of sect. Fimbriatum Williams and sect. Barbulatum Williams.

Cell Nucleus↗

Extraordinary haplotype diversity in haplodiploid inbreeders: phylogenetics and evolution of the bark beetle genus Coccotrypes.

Regular inbreeding by sib-mating is one of the most successful ecological strategies in the bark beetle family Scolytinae. Within this family, the many species (119) in Coccotrypes are found breeding in an exceptional variety of untraditional woody tissues different from bark and phloem. Species delineation by morphological criteria is extremely difficult, however, as in most other inbreeding groups of beetles, perhaps due to the unusual evolutionary dynamics characterizing sib-mating organisms. Hence, we here performed a phylogenetic analysis using molecular data in conjunction with morphological data to better understand morphological and ecological evolution in this sib-mating group. We used partial DNA sequences from the nuclear gene EF-alpha and the mitochondrial genes 12S and CO1 to elucidate patterns of morphological evolution, haplotype variation, and evolutionary pathways in resource use. Sequence variation was high among species and far above that expected at the species level (e.g., 19% for CO1 within Coccotrypes advena). The tendency for exhaustive sequence variation at deeper nodes resulted in ambiguous reconstructions of the deepest splits. However, all results suggested that species with the broadest diets were clustered in a single derived position-another piece of evidence against specialization as a derived evolutionary feature.

Animals↗

Genetic mapping of an insertional hydrocephalus-inducing mutation allelic to hy3.

The transgenic mouse line OVE459 carries a transgene-induced insertional mutation resulting in autosomal recessive congenital hydrocephalus. Homozygous transgenic animals experience ventricular dilation with perinatal onset and are noticeably smaller than hemizygous or non-transgenic littermates within a few days after birth. Fluorescence in situ hybridization (FISH) revealed that the transgene inserted in a single locus on mouse Chromosome (chr) 8, region D2-E1. Genetic crosses between hemizygous OVE459 mice and mice heterozygous for the spontaneous mutation hydrocephalus-3 (hy3) produced hydrocephalic offspring with a frequency of 22%, demonstrating that these two mutations are allelic. A genomic library was made by using DNA from homozygous OVE459 mice, and genomic DNA flanking the transgene insertion site was isolated and sequenced. A PCR polymorphism between C57BL/6 DNA and Mus spretus was used to map the location of the transgene insert to 1.06 cM +/- 0.75 proximal to D8Mit152 by using the Jackson Laboratory Backcross DNA Panel Mapping Resource. Furthermore, sequence analysis from a mouse bacterial artificial chromosome (BAC) clone, positive for unique markers on both sides of the transgene insertion site, demonstrated that the genomic DNAs flanking each side of the transgene insertion are physically separated by approximately 51 kb on the wild-type mouse chromosome.

Animals↗

A phylogenomic gene cluster resource: the Phylogenetically Inferred Groups (PhIGs) database.

BACKGROUND: We present here the PhIGs database, a phylogenomic resource for sequenced genomes. Although many methods exist for clustering gene families, very few attempt to create truly orthologous clusters sharing descent from a single ancestral gene across a range of evolutionary depths. Although these non-phylogenetic gene family clusters have been used broadly for gene annotation, errors are known to be introduced by the artifactual association of slowly evolving paralogs and lack of annotation for those more rapidly evolving. A full phylogenetic framework is necessary for accurate inference of function and for many studies that address pattern and mechanism of the evolution of the genome. The automated generation of evolutionary gene clusters, creation of gene trees, determination of orthology and paralogy relationships, and the correlation of this information with gene annotations, expression information, and genomic context is an important resource to the scientific community. DISCUSSION: The PhIGs database currently contains 23 completely sequenced genomes of fungi and metazoans, containing 409,653 genes that have been grouped into 42,645 gene clusters. Each gene cluster is built such that the gene sequence distances are consistent with the known organismal relationships and in so doing, maximizing the likelihood for the clusters to represent truly orthologous genes. The PhIGs website contains tools that allow the study of genes within their phylogenetic framework through keyword searches on annotations, such as GO and InterPro assignments, and sequence similarity searches by BLAST and HMM. In addition to displaying the evolutionary relationships of the genes in each cluster, the website also allows users to view the relative physical positions of homologous genes in specified sets of genomes. SUMMARY: Accurate analyses of genes and genomes can only be done within their full phylogenetic context. The PhIGs database and corresponding website http://phigs.org address this problem for the scientific community. Our goal is to expand the content as more genomes are sequenced and use this framework to incorporate more analyses.

Base Sequence↗

An ESTs description of the new Vanin gene family conserved from fly to human.

Circulation and tissue colonization are essential properties of lymphoid cells and involve major families of adhesion molecules (e.g. , integrin, selectin, mucin-like, and molecules from the immunoglobulin superfamily). The mouse Vanin-1 molecule was recently identified and found to be involved in the colonization of the thymus by hematopoietic precursor cells. Here we show based on computational analysis of EST sequence database resources that Vanin-1 belongs to a new family of related molecules present from drosophila to human. This family includes the amidase enzyme Biotinidase, and a central protein domain is shared between Vanin and Nitrilase families, suggesting that Vanin molecules might bear an enzymatic activity. Five of these molecules were new uncharacterized cDNA sequences only described as ESTs. The three human Vanin genes map to the same region of Chromosome 6q. The detailed results are consultable at the VANIN web page (http://tagc. univ-mrs.fr/pub/vanin/).

Amidohydrolases↗

The sequencing-based typing tool of dbMHC: typing highly polymorphic gene sequences.

The dbMHC resource (http://www.ncbi.nlm.nih.gov/mhc/sbt.cgi?cmd=main) at the National Center for Biotechnology Information (NCBI) has developed an online tool for evaluating the allelic composition of sequencing-based typing (SBT) results of cDNA or genomic sequences. Whether the samples are heterozygous, haploid or a combination of the two, they can be compared with two up-to-date databases of all known alleles of several human leukocyte antigen (HLA) and killer cell immunoglobulin-like receptor (KIR) loci. The results of the submission are returned as a table of potential allele hits, along with the respective base changes and an interactive sequence viewer for close examination of the alignment.

Algorithms↗

Uncovering new lineages in the Sunda pangolin (Manis javanica) with museum mitogenomics.

Accurately identifying evolutionarily significant units (ESUs) is crucial for conservation planning, especially for species like pangolins threatened by overhunting and habitat loss. ESUs help categorize different pangolin populations, aiding in understanding their genetic diversity and distribution, which is vital for targeted conservation efforts. This research generated mitochondrial genomes from historical museum specimens of Sunda pangolins (Manis javanica) from underrepresented locations, uncovering a new evolutionary lineage from the Mentawai Islands that diverged from Indochina and west Sundaland populations around 760 000 years ago. This population thereby represents a divergent ESU with a small distribution, important for conservation planning. The novel sequences provide resources for forensic labs tracing the origin of confiscated scales and shed light into the potential distribution of the 'mysterious pangolin'. Additionally, this research confirmed the presence of the two major M. javanica lineages in Java and extended the known distribution of the eastern clade to Bali and East Kalimantan. Our findings potentially suggest a recent bottleneck and postglacial expansion of pangolins across Indochina and west Sundaland. Further investigation with genomic and morphological evidence, contact area sampling and type sequencing will be required to evaluate the taxonomic status of different M. javanica lineages and M. culionensis.

Genomics↗

Spatial self-organization in a cyclic resource-species model.

Biological communities are remarkable in their ability to form cooperative ensembles that lead to coexistence through various types of niche partitioning, usually intimately tied to spatial structure. This is especially true in microbial settings where differential expression and regulation of genes allows members of a given species to alter their lifestyle so as to fill a functional role within the community. The resulting species interactions can involve feedback, as in the case of some bacterial consortia that participate in the cooperative degradation of a given resource in a succession of steps and in such a way that certain "later" species provide catalytic support for the primary degrader. We seek to capture the essential features of such spatially extended biological systems by introducing a lattice-based stochastic spatial model (interacting particle system) with cyclic local dynamics. Here, a given site progresses through a sequence of resource and species states in a prescribed order. Furthermore, this succession of states (at a site) is assumed to form a cyclic pattern due to a natural feedback mechanism. We explore conditions under which all the species are able to coexist and consider the extent to which this coexistence requires the development of spatio-temporal patterns, including spiral waves. This self-organization, if it occurs, results when synchronization of the dynamics at the microscopic level leads to macroscopic patterns. These patterns result in consumer-driven resource fluctuations that generate a form of spatio-temporal niche partitioning. As with most models of this complexity, we employ a mixture of mathematical analysis and simulations to develop an understanding of the resulting dynamics.

Adaptation, Biological↗

Cross-species sequence comparisons: a review of methods and available resources.

With the availability of whole-genome sequences for an increasing number of species, we are now faced with the challenge of decoding the information contained within these DNA sequences. Comparative analysis of DNA sequences from multiple species at varying evolutionary distances is a powerful approach for identifying coding and functional noncoding sequences, as well as sequences that are unique for a given organism. In this review, we outline the strategy for choosing DNA sequences from different species for comparative analyses and describe the methods used and the resources publicly available for these studies.

Animals↗

16S rRNA gene sequencing for bacterial pathogen identification in the clinical laboratory.

For many years, sequencing of the 16S ribosomal RNA (rRNA) gene has served as an important tool for determining phylogenetic relationships between bacteria. The features of this molecular target that make it a useful phylogenetic tool also make it useful for bacterial detection and identification in the clinical laboratory. Sequence analysis of the 16S rRNA gene is a powerful mechanism for identifying new pathogens in patients with suspected bacterial disease, and more recently this technology is being applied in the clinical laboratory for routine identification of bacterial isolates. Several studies have shown that sequence identification is useful for slow-growing, unusual, and fastidious bacteria as well as for bacteria that are poorly differentiated by conventional methods. The technical resources necessary for sequence identification are significant. This method requires reagents and instrumentation for amplification and sequencing, a database of known sequences, and software for sequence editing and database comparison. Commercial reagents are available, and laboratory-developed assays for amplification and sequencing have been reported. Likewise, there are an increasing number of commercial and public databases. Despite the availability of resources, sequence-based identification is still relatively expensive. The cost is significantly reduced only by the introduction of more automated methods. As the cost decreases, this technology is likely to be more widely applied in the clinical setting.

Bacteria↗

Group foraging sensitivity to predictable and unpredictable changes in food distribution: past experience or present circumstances?

The ideal free distribution theory (Fretwell & Lucas, 1970) predicts that the ratio of foragers at two patches will equal the ratio of food resources obtained at the two patches. The theory assumes that foragers have "perfect knowledge" of patch profitability and that patch choice maximizes fitness. How foragers assess patch profitability has been debated extensively. One assessment strategy may be the use of past experience with a patch. Under stable environmental conditions, this strategy enhances fitness. However, in a highly unpredictable environment, past experience may provide inaccurate information about current conditions. Thus, in a nonstable environment, a strategy that allows rapid adjustment to present circumstances may be more beneficial. Evidence for this type of strategy has been found in individual choice. In the present experiments, a flock of pigeons foraged at two patches for food items and demonstrated results similar to those found in individual choice. Experiment 1 utilized predictable and unpredictable sequences of resource ratios presented across days or within a single session. Current foraging decisions depended on past experience, but that dependence diminished when the current foraging environment became more unpredictable. Experiment 2 repeated Experiment I with a different flock of pigeons under more controlled circumstances in an indoor coop and produced similar results.

Animals↗

The Swiss-Prot protein knowledgebase and ExPASy: providing the plant community with high quality proteomic data and tools.

The Swiss-Prot protein knowledgebase provides manually annotated entries for all species, but concentrates on the annotation of entries from model organisms to ensure the presence of high quality annotation of representative members of all protein families. A specific Plant Protein Annotation Program (PPAP) was started to cope with the increasing amount of data produced by the complete sequencing of plant genomes. Its main goal is the annotation of proteins from the model plant organism Arabidopsis thaliana. In addition to bibliographic references, experimental results, computed features and sometimes even contradictory conclusions, direct links to specialized databases connect amino acid sequences with the current knowledge in plant sciences. As protein families and groups of plant-specific proteins are regularly reviewed to keep up with current scientific findings, we hope that the wealth of information of Arabidopsis origin accumulated in our knowledgebase, and the numerous software tools provided on the Expert Protein Analysis System (ExPASy) web site might help to identify and reveal the function of proteins originating from other plants. Recently, a single, centralized, authoritative resource for protein sequences and functional information, UniProt, was created by joining the information contained in Swiss-Prot, Translation of the EMBL nucleotide sequence (TrEMBL), and the Protein Information Resource-Protein Sequence Database (PIR-PSD). A rising problem is that an increasing number of nucleotide sequences are not being submitted to the public databases, and thus the proteins inferred from such sequences will have difficulties finding their way to the Swiss-Prot or TrEMBL databases.

Arabidopsis↗

Plant genomics: present state and a perspective on future developments.

The year 2001 may well be called the Year of the Human Genome. Less in the limelight, but equally exciting for plant scientists, is the rapid progress in plant genomics. With relatively modest resources, a lot has been achieved. The Arabidopsis genomic sequence (125 megabases [Mb]) is essentially finished, and rice sequencing is progressing rapidly. For many species, expressed sequence tag (EST) resources are plentiful, allowing broad inter-specific comparisons. At the same time, development of integrated physical-genetic maps for large-genome crop species is not progressing as rapidly as desired, while resources for the complete sequencing of these crops are not likely to become available. Some important plant genomes are so large that their complete sequencing may not be practical for many years. Significant plant genome research is concentrated in industry, and not freely available, creating some frustration in the academic community. Growing interest is anticipated in the development of metabolic profiling technologies, RNA profiling, proteomics and integrated systems approaches to plant biology.

Computational Biology↗

High-throughput development and characterization of a genomewide collection of gene-based single nucleotide polymorphism markers by chip-based matrix-assisted laser desorption/ionization time-of-flight mass spectrometry.

We describe here a system for the rapid identification, assay development, and characterization of gene-based single nucleotide polymorphisms (SNPs). This system couples informatics tools that mine candidate SNPs from public expressed sequence tag resources and automatically designs assay reagents with detection by a chip-based matrix-assisted laser desorption/ionization time-of-flight mass spectrometry platform. As a proof of concept of this system, a genomewide collection of reagents for 9,115 gene-based SNP genetic markers was rapidly developed and validated. These data provide preliminary insights into patterns of polymorphism in a genomewide collection of gene-based polymorphisms.

Alleles↗

Large-scale sequencing and the new animal phylogeny.

Although comparisons of gene sequences have revolutionised our understanding of the animal phylogenetic tree, it has become clear that, to avoid errors in tree reconstruction, a large number of genes from many species must be considered: too few genes and stochastic errors predominate, too few taxa and systematic errors appear. We argue here that, to gather many sequences from many taxa, the best use of resources is to sequence a small number of expressed sequence tags (1000-5000 per species) from as many taxa as possible. This approach counters both sources of error, gives the best hope of a well-resolved phylogeny of the animals and will act as a central resource for a carefully targeted genome sequencing programme.

Animals↗