Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequencing Resource”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

Isolation and analysis of candidate myeloid tumor suppressor genes from a commonly deleted segment of 7q22.

Monosomy 7 and deletions of 7q are recurring leukemia-associated cytogenetic abnormalities that correlate with adverse outcomes in children and adults. We describe a 2.52-Mb genomic DNA contig that spans a commonly deleted segment of chromosome band 7q22 identified in myeloid malignancies. This interval currently includes 14 genes, 19 predicted genes, and 5 predicted pseudogenes. We have extensively characterized the FBXL13, NAPE-PLD, and SVH genes as candidate myeloid tumor suppressors. FBXL13 encodes a novel F-box protein, SVHis a member of a gene family that contains Armadillo-like repeats, and NAPE-PLD encodes a phospholipase D-type phosphodiesterase. Analysis of a panel of leukemia specimens with monosomy 7 did not reveal mutations in these or in the candidate genes LRRC17, PRO1598, and SRPK2. This fully sequenced and annotated contig provides a resource for candidate myeloid tumor suppressor gene discovery.

Base Sequence↗

Rice mutant resources for gene discovery.

With the completion of genomic sequencing of rice, rice has been firmly established as a model organism for both basic and applied research. The next challenge is to uncover the functions of genes predicted by sequence analysis. Considering the amount of effort and the diversity of disciplines required for functional analyses, extensive international collaboration is needed for this next goal. The aims of this review are to summarize the current status of rice mutant resources, key tools for functional analysis of genes, and our perspectives on how to accelerate rice gene discovery through collaboration.

Databases, Genetic↗

Comparative EST analyses in plant systems.

Expressed sequence tag (EST) data are a major contributor to the known plant sequence space. Organization of the data into non-redundant clusters representing tentative unique genes provides snapshots of the gene repertoires of a species. This chapter reviews availability of sequences and sequence analysis results and describes several resources and tools that should facilitate broad-based utilization of EST data for gene structure annotation, gene discovery, and comparative genomics.

Alternative Splicing↗

A collection of 11 800 single-copy Ds transposon insertion lines in Arabidopsis.

More than 10 000 transposon-tagged lines were constructed by using the Activator (Ac)/Dissociation (Ds) system in order to collect insertional mutants as a useful resource for functional genomics of Arabidopsis. The flanking sequences of the Ds element in the 11 800 independent lines were determined by high-throughput analysis using a semi-automated method. The sequence data allowed us to map the unique insertion site on the Arabidopsis genome in each line. The Ds element of 7566 lines is inserted in or close to coding regions, potentially affecting the function of 5031 of 25 000 Arabidopsis genes. Half of the lines have Ds insertions on chromosome 1 (Chr. 1), in which donor lines have a donor site. In the other half, the Ds insertions are distributed throughout the other four chromosomes. The intrachromosomal distribution of Ds insertions varies with the donor lines. We found that there are hot spots for Ds transposition near the ends of every chromosome, and we found some statistical preference for Ds insertion targets at the nucleotide level. On the basis of systematic analysis of the Ds insertion sites in the 11 800 lines, we propose the use of Ds-tagged lines with a single insertion in annotated genes for systematic analysis of phenotypes (phenome analysis) in functional genomics. We have opened a searchable database of the insertion-site sequences and mutated genes (http://rarge.gsc.riken.go.jp/) and are depositing these lines in the RIKEN BioResource Center as available resources (http://www.brc.riken.go.jp/Eng/).

Arabidopsis↗

Genome-wide identification of nodule-specific transcripts in the model legume Medicago truncatula.

The Medicago truncatula expressed sequence tag (EST) database (Gene Index) contains over 140,000 sequences from 30 cDNA libraries. This resource offers the possibility of identifying previously uncharacterized genes and assessing the frequency and tissue specificity of their expression in silico. Because M. truncatula forms symbiotic root nodules, unlike Arabidopsis, this is a particularly important approach in investigating genes specific to nodule development and function in legumes. Our analyses have revealed 340 putative gene products, or tentative consensus sequences (TCs), expressed solely in root nodules. These TCs were represented by two to 379 ESTs. Of these TCs, 3% appear to encode novel proteins, 57% encode proteins with a weak similarity to the GenBank accessions, and 40% encode proteins with strong similarity to the known proteins. Nodule-specific TCs were grouped into nine categories based on the predicted function of their protein products. Besides previously characterized nodulins, other examples of highly abundant nodule-specific transcripts include plantacyanin, agglutinin, embryo-specific protein, and purine permease. Six nodule-specific TCs encode calmodulin-like proteins that possess a unique cleavable transit sequence potentially targeting the protein into the peribacteroid space. Surprisingly, 114 nodule-specific TCs encode small Cys cluster proteins with a cleavable transit peptide. To determine the validity of the in silico analysis, expression of 91 putative nodule-specific TCs was analyzed by macroarray and RNA-blot hybridizations. Nodule-enhanced expression was confirmed experimentally for the TCs composed of five or more ESTs, whereas the results for those TCs containing fewer ESTs were variable.

Agglutinins↗

Online Mendelian Inheritance in Man (OMIM).

Online Mendelian Inheritance In Man (OMIM) is a public database of bibliographic information about human genes and genetic disorders. Begun by Dr. Victor McKusick as the authoritative reference Mendelian Inheritance in Man, it is now distributed electronically by the National Center for Biotechnology Information (NCBI). Material in OMIM is derived from the biomedical literature and is written by Dr. McKusick and his colleagues at Johns Hopkins University and elsewhere. Each OMIM entry has a full text summary of a genetic phenotype and/or gene and has copious links to other genetic resources such as DNA and protein sequence, PubMed references, mutation databases, approved gene nomenclature, and more. In addition, NCBI's neighboring feature allows users to identify related articles from PubMed selected on the basis of key words in the OMIM entry. Through its many features, OMIM is increasingly becoming a major gateway for clinicians, students, and basic researchers to the ever-growing literature and resources of human genetics.

Alleles↗

Molecular biodiversity. Case study: Porifera (sponges).

Biological diversity--or biodiversity--is the term given to the variety of life on Earth and the natural patterns it forms. The biodiversity we see today is the fruit of billions of years of evolution, shaped by natural processes and, increasingly, by the influence of humans. It forms the web of life of which we are an integral part and upon which we so fully depend. The research on molecular biodiversity tries to lay the scientific foundation of a rational conservation policy that has its roots in various disciplines including systematics/taxonomy (species richness), present day ecology (diversity of ecological systems), and functional genetics (genetic diversity). The results of ongoing genome analyses (genome projects and expressed sequence tag projects) and the achievements of molecular evolution may allow us not only to quantitate the diversity of the present biota but also to extrapolate to their diversification in the future. A link between biodiversity and genomics/molecular evolution will create a platform which we hope may facilitate a sustainable management of organismic life and ensure its exploitation for human benefit. In the present review we outline possible strategies, using the Porifera (sponges) as a prominent example. On the basis of solid taxonomy and ecological data, the high value of this phylum for human application becomes obvious, especially with regard to the field of chemical ecology and the desire to find novel potential drugs for clinical use. In addition, the benefit of trying to make sense of molecular biodiversity using sponges as an example can be seen in the fact that the study of these animals, which are "living fossils", gives us a good insight into the history of our planet, especially with respect to the evolution of Metazoa.

Amino Acid Sequence↗

Molecular cloning of giant panda pituitary prolactin cDNA and its expression in Escherichia coli.

cDNA encoding pituitary (PRL) of giant panda was obtained using RT-PCR and expressed in E. coli. The results revealed that panda PRL cDNA encodes a precursor protein of 229 amino acids including a putative signal peptide of 30 amino acids and a mature protein of 199 residues with one potential N-glycosylation site. Sequence comparison indicated that panda PRL shares a high degree of identity to other known PRL sequences ranging from 98% with mink PRL to about 50% with rodent PRL. Six cysteine residues and 29 conserved residues distributed in four domains (PD1, PD2, PD3, and PD4) of PRL were observed. through multiple sequence alignment. Fourteen key residues of binding sites 1 and 2 involved in receptor binding are conserved in panda PRL. GST fused recombinant panda PRL protein was efficiently expressed with the form of insoluble inclusion bodies in E. coli BL21 transformed with a pGEX-4T-1 expression vector containing the DNA sequence encoding mature panda PRL. Western blot analysis indicated that GST-panda PRL recombinant protein could be recognized by antibody against human PRL. Our results would contribute to further elucidating the structural and functional characteristics of pituitary PRL and provide a basis for the production of recombinant panda prolactin for future use in the breeding of giant panda.

Amino Acid Sequence↗

Molecular characterization of the interferon-tau gene of the mithun (Bos frontalis).

The mithun (Bos frontalis) not only remains one of the most neglected ungulate species due to its remote range, but also has been identified as a vulnerable species due to its declining population. Augmenting its reproductive efficiency could be a strategy for reversing its population decline. Considering the importance of interferon-tau (IFNT) as a primary signal in establishing maternal recognition of pregnancy (MRP), the present study was undertaken to characterize the IFNT gene of the mithun. A 588 bp mithun IFNT (mitIFNT) gene was PCR amplified using genomic DNA as the template. Its nucleotide sequence comprised an entire open reading frame of 585 bp encoding a 195 amino acid pre-protein. In nucleotide sequence, the mitIFNT gene was more than 85% similar to the homologous genes of domestic and wild ruminant species characterized to date. However, phylogenetic analysis placed mitIFNT into a clade containing IFNT of the red deer, but not IFNTs of cow, sheep, or goats, or other wild ruminant species. Our characterization of mitIFNT represents the first complete sequence of any gene from the mithun.

Amino Acid Sequence↗

RefSeq and LocusLink: NCBI gene-centered resources.

Thousands of genes have been painstakingly identified and characterized a few genes at a time. Many thousands more are being predicted by large scale cDNA and genomic sequencing projects, with levels of evidence ranging from supporting mRNA sequence and comparative genomics to computing ab initio models. This, coupled with the burgeoning scientific literature, makes it critical to have a comprehensive directory for genes and reference sequences for key genomes. The NCBI provides two resources, LocusLink and RefSeq, to meet these needs. LocusLink organizes information around genes to generate a central hub for accessing gene-specific information for fruit fly, human, mouse, rat and zebrafish. RefSeq provides reference sequence standards for genomes, transcripts and proteins; human, mouse and rat mRNA RefSeqs, and their corresponding proteins, are discussed here. Together, RefSeq and LocusLink provide a non-redundant view of genes and other loci to support research on genes and gene families, variation, gene expression and genome annotation. Additional information about LocusLink and RefSeq is available at http://www.ncbi.nlm.nih.gov/LocusLink/.

Animals↗

Optimizing radiologic workup: an artificial intelligence approach.

The increasing complexity of diagnostic imaging is presenting an ever expanding variety of radiologic test options to clinicians. As a result, it is becoming more difficult for referring physicians to select an appropriate sequence of tests. The current economic pressures on medicine make it particularly important that resources be used judiciously. Radiologic workup often involves a sequence of tests that lead from presenting signs and symptoms to a definitive diagnosis or intervention. This sequence ideally begins with simple, inexpensive, safe, non-invasive tests and progresses to more complex, expensive, and hazardous tests only if the simpler tests are insufficient to establish a diagnosis. DxCON is a developmental artificial intelligence-based computer system that gives advice to physicians about the optimum sequencing of radiologic tests. DxCON evaluates basic clinical information and a physician's proposed workup plan. The system then creates an analysis of the strengths and weaknesses of his plan. The domain chosen to explore computer-based workup advice is the radiologic workup of obstructive jaundice.

Cholestasis↗

Subtree power analysis and species selection for comparative genomics.

Sequence comparison across multiple organisms aids in the detection of regions under selection. However, resource limitations require a prioritization of genomes to be sequenced. This prioritization should be grounded in two considerations: the lineal scope encompassing the biological phenomena of interest, and the optimal species within that scope for detecting functional elements. We introduce a statistical framework for optimal species subset selection, based on maximizing power to detect conserved sites. Analysis of a phylogenetic star topology shows theoretically that the optimal species subset is not in general the most evolutionarily diverged subset. We then demonstrate this finding empirically in a study of vertebrate species. Our results suggest that marsupials are prime sequencing candidates.

Animals↗

Analysis of bovine mammary gland EST and functional annotation of the Bos taurus gene index.

Functional genomic studies of the mammary gland require an appropriate collection of cDNA sequences to assess gene expression patterns from the different developmental and operational states of underlying cell types. To better capture the range of gene expression, a normalized cDNA library was constructed from pooled bovine mammary tissues, and 23,202 expressed sequence tags (EST) were produced and deposited into GenBank. Assembly of these EST with sequences in the Bos taurus Gene Index (BtGI) helped to form 5751 of the current 23,883 tentative consensus (TC) sequences. The majority (87%) of these 5751 assemblies contained only one to three mammary-derived EST. In contrast, 18% of the mammary EST assembled with TC sequences corresponding to 12 genes. These results suggest library normalization was only partially effective, because the reduction in EST for genes abundantly transcribed during lactation could be attributed to pooling. For better assessment of novel content in the mammary library and to add to existing annotation of all bovine sequence elements, gene ontology assignments, and comparative sequence analyses against human genome sequence, human and rodent gene indices, and an index of orthologous alignments of genes across eukaryotes (TOGA) were performed, and results were added to existing BtGI annotation. Over 35,000 of the bovine elements significantly matched human genome sequence, and the positions of some alignments (3%) were unique relative to those using human expressed sequences. Because 3445 TC sequences had no significant match with any data set, mammary-derived cDNA clones representing 23 of these elements were analyzed further for expression and novelty. Only one clone met criteria suggesting the corresponding gene was a divergent ortholog or expressed sequence unique to cattle. These results demonstrate that bovine sequence expression data serve as a resource for characterizing mammalian transcriptomes and identifying those genes potentially unique to ruminants.

Animals↗

The TIGR Maize Database.

Maize is a staple crop of the grass family and also an excellent model for plant genetics. Owing to the large size and repetitiveness of its genome, we previously investigated two approaches to accelerate gene discovery and genome analysis in maize: methylation filtration and high C(0)t selection. These techniques allow the construction of gene-enriched genomic libraries by minimizing repeat sequences due to either their methylation status or their copy number, yielding a 7-fold enrichment in genic sequences relative to a random genomic library. Approximately 900,000 gene-enriched reads from maize were generated and clustered into Assembled Zea mays (AZM) sequences. Here we report the current AZM release, which consists of approximately 298 Mb representing 243,807 sequence assemblies and singletons. In order to provide a repository of publicly available maize genomic sequences, we have created the TIGR Maize Database (http://maize.tigr.org). In this resource, we have assembled and annotated the AZMs and used available sequenced markers to anchor AZMs to maize chromosomes. We have constructed a maize repeat database and generated draft sequence assemblies of 287 maize bacterial artificial chromosome (BAC) clone sequences, which we annotated along with 172 additional publicly available BAC clones. All sequences, assemblies and annotations are available at the project website via web interfaces and FTP downloads.

Chromosome Mapping↗

ELM server: A new resource for investigating short functional sites in modular eukaryotic proteins.

Multidomain proteins predominate in eukaryotic proteomes. Individual functions assigned to different sequence segments combine to create a complex function for the whole protein. While on-line resources are available for revealing globular domains in sequences, there has hitherto been no comprehensive collection of small functional sites/motifs comparable to the globular domain resources, yet these are as important for the function of multidomain proteins. Short linear peptide motifs are used for cell compartment targeting, protein-protein interaction, regulation by phosphorylation, acetylation, glycosylation and a host of other post-translational modifications. ELM, the Eukaryotic Linear Motif server at http://elm.eu.org/, is a new bioinformatics resource for investigating candidate short non-globular functional motifs in eukaryotic proteins, aiming to fill the void in bioinformatics tools. Sequence comparisons with short motifs are difficult to evaluate because the usual significance assessments are inappropriate. Therefore the server is implemented with several logical filters to eliminate false positives. Current filters are for cell compartment, globular domain clash and taxonomic range. In favourable cases, the filters can reduce the number of retained matches by an order of magnitude or more.

Amino Acid Motifs↗

Overview of the yeast genome.

The collaboration of more than 600 scientists from over 100 laboratories to sequence the Saccharomyces cerevisiae genome was the largest decentralised experiment in modern molecular biology and resulted in a unique data resource representing the first complete set of genes from a eukaryotic organism. 12 million bases were sequenced in a truly international effort involving European, US, Canadian and Japanese laboratories. While the yeast genome represents only a small fraction of the information in today's public sequence databases, the complete, ordered and non-redundant sequence provides an invaluable resource for the detailed analysis of cellular gene function and genome architecture. In terms of throughput, completeness and information content, yeast has always been the lead eukaryotic organism in genomics; it is still the largest genome to be completely sequenced.

Chromosome Mapping↗

GO-Diff: mining functional differentiation between EST-based transcriptomes.

BACKGROUND: Large-scale sequencing efforts produced millions of Expressed Sequence Tags (ESTs) collectively representing differentiated biochemical and functional states. Analysis of these EST libraries reveals differential gene expressions, and therefore EST data sets constitute valuable resources for comparative transcriptomics. To translate differentially expressed genes into a better understanding of the underlying biological phenomena, existing microarray analysis approaches usually involve the integration of gene expression with Gene Ontology (GO) databases to derive comparable functional profiles. However, methods are not available yet to process EST-derived transcription maps to enable GO-based global functional profiling for comparative transcriptomics in a high throughput manner. RESULTS: Here we present GO-Diff, a GO-based functional profiling approach towards high throughput EST-based gene expression analysis and comparative transcriptomics. Utilizing holistic gene expression information, the software converts EST frequencies into EST Coverage Ratios of GO Terms. The ratios are then tested for statistical significances to uncover differentially represented GO terms between the compared transcriptomes, and functional differences are thus inferred. We demonstrated the validity and the utility of this software by identifying differentially represented GO terms in three application cases: intra-species comparison; meta-analysis to test a specific hypothesis; inter-species comparison. GO-Diff findings were consistent with previous knowledge and provided new clues for further discoveries. A comprehensive test on the GO-Diff results using series of comparisons between EST libraries of human and mouse tissues showed acceptable levels of consistency: 61% for human-human; 69% for mouse-mouse; 47% for human-mouse. CONCLUSION: GO-Diff is the first software integrating EST profiles with GO knowledge databases to mine functional differentiation between biological systems, e.g. tissues of the same species or the same tissue cross species. With rapid accumulation of EST resources in the public domain and expanding sequencing effort in individual laboratories, GO-Diff is useful as a screening tool before undertaking serious expression studies.

Animals↗