Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequencing Resource”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Investigating semantic similarity measures across the Gene Ontology: the relationship between sequence and annotation.

MOTIVATION: Many bioinformatics data resources not only hold data in the form of sequences, but also as annotation. In the majority of cases, annotation is written as scientific natural language: this is suitable for humans, but not particularly useful for machine processing. Ontologies offer a mechanism by which knowledge can be represented in a form capable of such processing. In this paper we investigate the use of ontological annotation to measure the similarities in knowledge content or 'semantic similarity' between entries in a data resource. These allow a bioinformatician to perform a similarity measure over annotation in an analogous manner to those performed over sequences. A measure of semantic similarity for the knowledge component of bioinformatics resources should afford a biologist a new tool in their repertoire of analyses. RESULTS: We present the results from experiments that investigate the validity of using semantic similarity by comparison with sequence similarity. We show a simple extension that enables a semantic search of the knowledge held within sequence databases. AVAILABILITY: Software available from http://www.russet.org.uk.

Artificial Intelligence↗

Molecular breeding of tomato: Advances and challenges.

The modern cultivated tomato (Solanum lycopersicum) was domesticated from Solanum pimpinellifolium native to the Andes Mountains of South America through a "two-step domestication" process. It was introduced to Europe in the 16th century and later widely cultivated worldwide. Since the late 19th century, breeders, guided by modern genetics, breeding science, and statistical theory, have improved tomatoes into an important fruit and vegetable crop that serves both fresh consumption and processing needs, satisfying diverse consumer demands. Over the past three decades, advancements in modern crop molecular breeding technologies, represented by molecular marker technology, genome sequencing, and genome editing, have significantly transformed tomato breeding paradigms. This article reviews the research progress in the field of tomato molecular breeding, encompassing genome sequencing of germplasm resources, the identification of functional genes for agronomic traits, and the development of key molecular breeding technologies. Based on these advancements, we also discuss the major challenges and perspectives in this field.

Solanum lycopersicum↗

PlantSat: a specialized database for plant satellite repeats.

MOTIVATION: Tandemly organized repetitive sequences (satellite DNA) are widespread in complex eukaryotic genomes. In plants, satellite repeats often represent a substantial part of nuclear DNA but only a little is known about the molecular mechanisms of their amplification and their possible role(s) in genome evolution and function. Unfortunately, addressing these questions via characterization of general sequence properties of known satellite repeats has been hindered by a difficulty in obtaining a complete and unbiased set of sequence data for this analysis. This is mainly due to the presence of multiple entries of homologous sequences and of single entries that contain more than one repeated unit (monomer) in the public databases. RESULTS: We have established a computer database specialized for plant satellite repeats (PlantSat) that integrates sequence data available from various resources with supplementary information including repeat consensus sequences, abundances, and chromosomal localizations. The sequences are stored as individual repeat monomers grouped into families, which simplifies their computer analysis and makes it more accurate. Using this feature, we have performed a basic sequence analysis of the whole set of plant satellite repeats with respect to their monomer length and nucleotide composition. The analysis revealed several preferred length ranges of the monomers (approximately 165 bp and its multiples) and an over-representation of the AA/TT dinucleotide in the repeats. We have also detected an enrichment of satellite DNA sequences for the motif CAAAA that is supposed to be involved in breakage-reunion of repeated sequences.

Computational Biology↗

A graph based algorithm for generating EST consensus sequences.

MOTIVATION: EST sequences constitute an abundant, yet error prone resource for computational biology. Expressed sequences are important in gene discovery and identification, and they are also crucial for the discovery and classification of alternative splicing. An important challenge when processing EST sequences is the reconstruction of mRNA by assembling EST clusters into consensus sequences. RESULTS: In contrast to the more established assembly tools, we propose an algorithm that constructs a graph over sequence fragments of fixed size, and produces consensus sequences as traversals of this graph. We provide a tool implementing this algorithm, and perform an experiment where the consensus sequences produced by our implementation, as well as by currently available tools, are compared to mRNA. The results show that our proposed algorithm in a majority of the cases produces consensus of higher quality than the established sequence assemblers and at a competitive speed. AVAILABILITY: The source code for the implementation is available under a GPL license from http://www.ii.uib.no/~ketil/bioinformatics/ CONTACT: ketil@ii.uib.no.

Algorithms↗

Genomic resources for chicken.

The recent sequencing and draft assembly of a chicken genome has provided biologists with an invaluable research tool that complements a growing list of additional avian genomic resources. For many researchers, finding and using these resources is challenging, because information is presented through an increasing number of Web sites and browser navigation frequently requires specific knowledge and expertise. This primer provides an overview of online genomic resources for the chicken, including the Ensembl, UCSC, and NCBI annotated chicken genome browsers; expressed sequence tag and in situ hybridization databases; and sources for microarrays, cDNAs, and bacterial artificial chromosomes (BACs). Several short tutorials oriented toward the biologist with limited bioinformatics skills outline how to retrieve several types of commonly needed information and reagents.

Animals↗

Transcription mapping of the 5q- syndrome critical region: cloning of two novel genes and sequencing, expression, and mapping of a further six novel cDNAs.

The 5q- syndrome is a myelodysplastic syndrome with the 5q deletion ¿del(5q) as the sole karyotypic abnormality. We are using the expressed sequence tag (EST) resource as our primary approach to identifying novel candidate genes for the 5q- syndrome. Seventeen ESTs were identified from the Human Gene Map at the National Center for Biotechnology Information that had no significant homology to any known genes and were assigned between DNA markers D5S413 and D5S487, flanking the critical region of the 5q- syndrome at 5q31-q32. Eleven of the 17 cDNAs from which the ESTs were derived (65%) were shown to map to the critical region of the 5q- syndrome by gene dosage analysis and were then sublocalized by PCR screening to a YAC contig encompassing the critical region. Eight of the 11 cDNA clones, upon full sequencing, had no significant homology to any known genes. Each of the 8 cDNA clones was shown to be expressed in human bone marrow. The complete coding sequence was obtained for 2 of the novel genes, termed C5orf3 and C5orf4. The 2.6-kb transcript of C5orf3 encodes a putative 505-amino-acid protein and contains an ATP/GTP-binding site motif A (P loop), suggesting that this novel gene encodes an ATP- or a GTP-binding protein. The novel gene C5orf4 has a transcript of 3.1 kb, encoding a putative 144-amino-acid protein. We describe the cloning of 2 novel human genes and the sequencing, expression patterns, and mapping to the critical region of the 5q- syndrome of a further 6 novel cDNA clones. Genomic localization and expression patterns would suggest that the 8 novel cDNAs described in this report represent potential candidate genes for the 5q- syndrome.

Amino Acid Sequence↗

The human transcript database: a catalogue of full length cDNA inserts.

SUMMARY: Full length cDNA sequences are an important resource for the research community but are currently intermingled with other sequences. We have identified the human full length insert cDNA sequences in GenBank and placed them in a single location, the Human Transcript Database. AVAILIBILITY: The Human Transcript Database is available at http://www.hgsc.bcm.tms.edu/HTDB/. CONTACT: John Bouck: jbouck@bcm.tmc.edu

DNA Transposable Elements↗

SAGE of the developing wheat caryopsis.

Understanding the development of the cereal caryopsis holds the future for metabolic engineering in the interests of enhancing global food production. We have developed a Serial Analysis of Gene Expression (SAGE) data platform to investigate the developing wheat (Triticum aestivum) caryopsis. LongSAGE libraries have been constructed at five time-points post-anthesis to coincide with key processes in caryopsis development. More than 90,000 LongSAGE tags have been sequenced generating 29,261 unique tag sequences across all five libraries. Tag abundance, generated from cumulative tag counts, provides insight into the redundancy and diversity of each library. Annotation of the 500 most abundant tags spanning development highlights the array of functional groups being expressed. The relative frequency of these more abundant transcripts allows quantitative analysis of patterns of expression during grain development. We have identified activities of cellular proliferation/differentiation, the accumulation of storage proteins and starch biosynthesis. The abundance of calcium-dependent protein kinases indicate their importance in signalling across development. Acquisition of a broad array of defence coincides with storage accumulation and is dominated by inhibitors of amylase activity. Differential expression profiles of abundant tags from each library reveal the coordinated expression of genes responsible for the cellular events constituting caryopsis development. This SAGE platform has also provided a resource of novel sequence and expression information including the identification of potentially useful promoter activities. Further investigations into both the abundant and low expressing transcripts will provide greater insight into wheat caryopsis development and assist in wheat improvement programmes.

Gene Expression Profiling↗

CR-EST: a resource for crop ESTs.

The crop expressed sequence tag database, CR-EST (http://pgrc.ipk-gatersleben.de/cr-est/), is a publicly available online resource providing access to sequence, classification, clustering and annotation data of crop EST projects. CR-EST currently holds more than 200,000 sequences derived from 41 cDNA libraries of four species: barley, wheat, pea and potato. The barley section comprises approximately one-third of all publicly available ESTs. CR-EST deploys an automatic EST preparation pipeline that includes the identification of chimeric clones in order to transparently display the data quality. Sequences are clustered in species-specific projects to currently generate a non-redundant set of approximately 22,600 consensus sequences and approximately 17,200 singletons, which form the basis of the provided set of unigenes. A web application allows the user to compute BLAST alignments of query sequences against the CR-EST database, query data from Gene Ontology and metabolic pathway annotations and query sequence similarities from stored BLAST results. CR-EST also features interactive JAVA-based tools, allowing the visualization of open reading frames and the explorative analysis of Gene Ontology mappings applied to ESTs.

Crops, Agricultural↗

The genetics of psoriasis 2001: the odyssey continues.

Accumulating evidence indicates that psoriasis is a multifactorial disorder caused by the concerted action of multiple disease genes in a single individual, triggered by environmental factors. Some of these genes control the severity of multiple diseases by regulating inflammation and immunity (severity genes), whereas others are unique to psoriasis. Various combinations of these genes can occur even within a single family, accounting in large measure for the many clinical manifestations of psoriasis. The disease-causing variants (alleles) of these genes probably arose early in the history of modern humans. As a result, psoriasis disease alleles are common in the general population, have a worldwide distribution, and often share the same ancestral chromosome with neutral alleles at adjacent loci. This phenomenon, called linkage disequilibrium, explains why psoriasis is strongly associated with HLA-Cw6 worldwide, although HLA-Cw6 is unlikely to be the disease allele. Many unaffected individuals carry 1 or more disease alleles, but lack other genetic and/or environmental factors necessary to produce disease. This explains why psoriasis develops in only about 10% of HLA-Cw6-positive individuals, and why genome-wide linkage scans for psoriasis and other multifactorial genetic disorders have not been uniformly successful. The Human Genome Project is rapidly generating a catalog of human DNA sequence variations. This resource has already allowed precise linkage disequilibrium mapping of the major histocompatibility complex psoriasis gene to just beyond HLA-C, toward HLA-A. This gene is likely to be identified soon. Further development and use of linkage disequilibrium resources will provide a powerful tool for the identification of the remaining psoriasis genes.

Genetic Heterogeneity↗

Gene-representing cDNA clusters defined by hybridization of 57,419 clones from infant brain libraries with short oligonucleotide probes.

Diverse biochemical and computational procedures and facilities have been developed to hybridize thousands of DNA clones with short oligonucleotide probes and subsequently to extract valuable genetic information. This technology has been applied to 73,536 cDNA clones from infant brain libraries. By a mutual comparison of 57,419 samples that were successfully scored by 200-320 probes, 19,726 genes have been identified and sorted by their expression levels. The data indicate that an additional 20,000 or more genes may be expressed in the infant brain. Representative clones of the found genes create a valuable resource for complete sequencing and functional studies of many novel genes. These results demonstrate the unique capacity of hybridization technology to identify weakly transcribed genes and to study gene networks involved in organismal development, aging, or tumorigenesis by monitoring the expression of every gene in related tissues, whether known or still undiscovered.

Base Sequence↗

An integrated gene and SSLP BAC map framework of mouse chromosome 11.

Physical maps are important resources both in sequencing and in functional analyses of large genomes. Global contig-building approaches are regarded to be more efficient relative to the cumulative outcome of scattered and more localized physical mapping studies accompanying positional cloning. This work is part of an effort to assemble a complete physical map of mouse chromosome 11 in which selection of clones containing specific genetic markers from genomic libraries is the first step in the process. Using a previously developed strategy, we identified 361 bacterial artificial chromosomes (BACs) containing 88 gene markers. Since the linkage positions of markers chosen for these studies are known, the BAC framework obtained is anchored to the genetic map and represents about 13% of the length of the entire chromosome. Together with similar assignments of BACs generated previously using D11Mit markers (Cai et al., 1988, Genomics, 54: 387-397), 36-40% of the chromosome 11 is now assembled into contigs, and these contigs correlate through 51 clones carrying both gene and simple sequence length polymorphism markers.

Animals↗

Characterization of 4AOHW cell line panel including new data for the 10IHW panel.

There will be a continuing need for well characterized panels of EBV-transformed lymphoblastoid cell lines. Selection of the 4AOH panel was based on prior MHC typing and was intended to ensure representation of ancestral haplotypes from various racial groups. Cells from nonhuman primates, bone marrow donor-recipient pairs, and patients with IDDM were included. Selected cells from the 10IHW were included to enable further characterization. Cells were distributed to participants in the 4AOHW and were typed at multiple loci by a variety of procedures. Non-HLA genes such as TNF were included. Since the cells were distributed "blind" with hidden replicates, it was possible to evaluate the quality of the typing data. An approach to data management is described. The best current estimates of the typing of these cells are presented. The panel will be useful since it provides standards for most alleles at most loci. Since the cells are so well characterized, they represent a useful resource for MHC sequencing and for the evaluation of new typing procedures.

Alleles↗

Genetic relatedness of human DNA polymerase beta and terminal deoxynucleotidyltransferase.

The Protein Identification Resource (PIR) protein sequence data bank was searched for sequence similarity between known proteins and human DNA polymerase beta (Pol beta) or human terminal deoxynucleotidyltransferase (TdT). Pol beta and TdT were found to exhibit amino acid sequence similarity only with each other and not with any other of the 4750 entries in release 12.0 of the PIR data bank. Optimal amino acid sequence alignment of the entire 39-kDa Pol beta polypeptide with the C-terminal two thirds of TdT revealed 24% identical aa residues and 21% conservative aa substitutions. The Monte Carlo score of 12.6 for the entire aligned sequences indicates highly significant aa sequence homology. The hydropathicity profiles of the aligned aa sequences were remarkably similar throughout, suggesting structural similarity of the polypeptides. The most significant regions of homology are aa residues 39-224 and 311-333 of Pol beta vs. aa residues 191-374 and 484-506 of TdT. In addition, weaker homology was seen between a large portion of the 'nonessential' N-terminal end of TdT (aa residues 33-130) and the first region of strong homology between the two proteins (aa residues 31-128 of Pol beta and aa residues 183-280 of TdT), suggestive of genetic duplication within the ancestral gene. On the basis of nucleotide differences between conserved regions of Pol beta and TdT genes (aligned according to optimally aligned aa sequences) it was estimated that Pol beta and TdT diverged on the order of 250 million years ago, corresponding roughly to a time before radiation of mammals and birds.

Amino Acid Sequence↗

Genomic characteristics and tracing analysis of an acute gastroenteritis outbreak associated with rotavirus C in a boarding high school.

BACKGROUND: Rotaviruses are major pathogens of childhood acute gastroenteritis, dominated by rotavirus A (RVA). Outbreaks caused by human rotavirus C (RVC) are rarely reported, and relevant genomic data remain scarce. This genomic investigation of an RVC outbreak improves our understanding of viral diversity and transmission dynamics. METHODS: We performed epidemiological surveys, nucleic acid testing and whole-genome sequencing (WGS) on specimens from a 2025 RVC-associated gastroenteritis outbreak at a Chinese boarding high school. Sequence alignment, phylogenetic and molecular tracing analyses were conducted to explore RVC evolution via point mutation, segment reassortment and genomic recombination. RESULTS: This typical point-source campus outbreak was linked to an indoor student gathering matching the incubation period of RVC. Thirteen RVC FX strains were recovered from 11 rectal swabs and two vomitus samples. Their viral protein (VP) 4 and VP7 sequences shared high homology with Russian reference strains, carrying distinct amino acid variations. No segment reassortment or recombination was detected in VP4/VP7 genes. CONCLUSIONS: Dense, closed campus settings facilitate RVC clustered transmission. Limitations included absent screening of asymptomatic canteen staff. Rapid nucleic acid testing enabled timely pathogen identification for outbreak control. Greater attention should be paid to the public health risk of RVC. These whole-genome sequencing data enrich resources for studying RVC evolution and vaccine development.

Acute gastroenteritis outbreak↗

Molecular and functional diversity of maize.

Over the past 10,000 years, man has used the rich genetic diversity of the maize genome as the raw material for domestication and subsequent crop improvement. Recent research efforts have made tremendous strides toward characterizing this diversity: structural diversity appears to be largely mediated by helitron transposable elements, patterns of diversity are yielding insights into the number and type of genes involved in maize domestication and improvement, and functional diversity experiments are leading to allele mining for future crop improvement. The development of genome sequence and germplasm resources are likely to further accelerate this progress.

Genetic Variation↗

A real-time and dynamic biological information retrieval and analysis system (BIRAS).

The aim of this study is to design a biological information retrieval and analysis system (BIRAS) based on the Internet. Using the specific network protocol, BIRAS system could send and receive information from the Entrez search and retrieval system maintained by National Center for Biotechnology Information (NCBI) in USA. The literatures, nucleotide sequence, protein sequences, and other resources according to the user-defined term could then be retrieved and sent to the user by pop up message or by E-mail informing automatically using BIRAS system. All the information retrieving and analyzing processes are done in real-time. As a robust system for intelligently and dynamically retrieving and analyzing on the user-defined information, it is believed that BIRAS would be extensively used to retrieve specific information from large amount of biological databases in now days. The program is available on request from the corresponding author.

Animals↗

Pan-genome based on chromosome sequences of wild and cultivated Agaricus bisporus.

Agaricus bisporus, one of the most widely cultivated mushrooms around the world, plays an important role in economy and agriculture. In this study, by employing long-reads generated by PacBio and Nanopore sequencing, we assembled six novel high-quality genomes (of which three are telomere-to-telomere assemblies) with sizes 29.6 ~ 30.8 Mb and N50 lengths of 2.5 ~ 2.6 Mb. Combined with public genome data of nine strains, we successfully established a pan-genome of A. bisporus, comprising a total of 14,626 clusters of protein coding genes, of which 50.70%, 7.45%, 24.74% and 17.01% are defined as core, soft core, dispensable, and private clusters, respectively. A total of 5,646 non- redundant structural variants (SVs) were identified among wild and cultivated strains and the genes associated with SV were mapped. This work provides valuable whole-genome sequences and genomic resources across wild and cultivated strains of the most widely cultivated mushroom species for functional analyses of genomes.

Agaricus↗