Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “CATALOGING”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

A comprehensive catalog of human KRAB-associated zinc finger genes: insights into the evolutionary history of a large family of transcriptional repressors.

Krüppel-type zinc finger (ZNF) motifs are prevalent components of transcription factor proteins in all eukaryotes. KRAB-ZNF proteins, in which a potent repressor domain is attached to a tandem array of DNA-binding zinc-finger motifs, are specific to tetrapod vertebrates and represent the largest class of ZNF proteins in mammals. To define the full repertoire of human KRAB-ZNF proteins, we searched the genome sequence for key motifs and then constructed and manually curated gene models incorporating those sequences. The resulting gene catalog contains 423 KRAB-ZNF protein-coding loci, yielding alternative transcripts that altogether predict at least 742 structurally distinct proteins. Active rounds of segmental duplication, involving single genes or larger regions and including both tandem and distributed duplication events, have driven the expansion of this mammalian gene family. Comparisons between the human genes and ZNF loci mined from the draft mouse, dog, and chimpanzee genomes not only identified 103 KRAB-ZNF genes that are conserved in mammals but also highlighted a substantial level of lineage-specific change; at least 136 KRAB-ZNF coding genes are primate specific, including many recent duplicates. KRAB-ZNF genes are widely expressed and clustered genes are typically not coregulated, indicating that paralogs have evolved to fill roles in many different biological processes. To facilitate further study, we have developed a Web-based public resource with access to gene models, sequences, and other data, including visualization tools to provide genomic context and interaction with other public data sets.

Computational Biology↗

Toward a catalog of human genes and proteins: sequencing and analysis of 500 novel complete protein coding human cDNAs.

With the complete human genomic sequence being unraveled, the focus will shift to gene identification and to the functional analysis of gene products. The generation of a set of cDNAs, both sequences and physical clones, which contains the complete and noninterrupted protein coding regions of all human genes will provide the indispensable tools for the systematic and comprehensive analysis of protein function to eventually understand the molecular basis of man. Here we report the sequencing and analysis of 500 novel human cDNAs containing the complete protein coding frame. Assignment to functional categories was possible for 52% (259) of the encoded proteins, the remaining fraction having no similarities with known proteins. By aligning the cDNA sequences with the sequences of the finished chromosomes 21 and 22 we identified a number of genes that either had been completely missed in the analysis of the genomic sequences or had been wrongly predicted. Three of these genes appear to be present in several copies. We conclude that full-length cDNA sequencing continues to be crucial also for the accurate identification of genes. The set of 500 novel cDNAs, and another 1000 full-coding cDNAs of known transcripts we have identified, adds up to cDNA representations covering 2%--5 % of all human genes. We thus substantially contribute to the generation of a gene catalog, consisting of both full-coding cDNA sequences and clones, which should be made freely available and will become an invaluable tool for detailed functional studies.

3' Untranslated Regions↗

CoreGenes: a computational tool for identifying and cataloging "core" genes in a set of small genomes.

BACKGROUND: Improvements in DNA sequencing technology and methodology have led to the rapid expansion of databases comprising DNA sequence, gene and genome data. Lower operational costs and heightened interest resulting from initial intriguing novel discoveries from genomics are also contributing to the accumulation of these data sets. A major challenge is to analyze and to mine data from these databases, especially whole genomes. There is a need for computational tools that look globally at genomes for data mining. RESULTS: CoreGenes is a global JAVA-based interactive data mining tool that identifies and catalogs a "core" set of genes from two to five small whole genomes simultaneously. CoreGenes performs hierarchical and iterative BLASTP analyses using one genome as a reference and another as a query. Subsequent query genomes are compared against each newly generated "consensus." These iterations lead to a matrix comprising related genes from this set of genomes, e. g., viruses, mitochondria and chloroplasts. Currently the software is limited to small genomes on the order of 330 kilobases or less. CONCLUSION: A computational tool CoreGenes has been developed to analyze small whole genomes globally. BLAST score-related and putatively essential "core" gene data are displayed as a table with links to GenBank for further data on the genes of interest. This web resource is available at http://pumpkins.ib3.gmu.edu:8080/CoreGenes or http://www.bif.atcc.org/CoreGenes.

Algorithms↗

Incidence of "quasi-ditags" in catalogs generated by Serial Analysis of Gene Expression (SAGE).

BACKGROUND: Serial Analysis of Gene Expression (SAGE) is a functional genomic technique that quantitatively analyzes the cellular transcriptome. The analysis of SAGE libraries relies on the identification of ditags from sequencing files; however, the software used to examine SAGE libraries cannot distinguish between authentic versus false ditags ("quasi-ditags"). RESULTS: We provide examples of quasi-ditags that originate from cloning and sequencing artifacts (i.e. genomic contamination or random combinations of nucleotides) that are included in SAGE libraries. We have employed a mathematical model to predict the frequency of quasi-ditags in random nucleotide sequences, and our data show that clones containing less than or equal to 2 ditags (which include chromosomal cloning artifacts) should be excluded from the analysis of SAGE catalogs. CONCLUSIONS: Cloning and sequencing artifacts contaminating SAGE libraries could be eliminated using simple pre-screening procedure to increase the reliability of the data.

Animals↗

Bacterial genotyping by 16S rRNA mass cataloging.

BACKGROUND: It has recently been demonstrated that organism identifications can be recovered from mass spectra using various methods including base-specific fragmentation of nucleic acids. Because mass spectrometry is extremely rapid and widely available such techniques offer significant advantages in some applications. A key element in favor of mass spectrometric analysis of RNA fragmentation patterns is that a reference database for analysis of the results can be generated from sequence information. In contrast to hybridization approaches, the genetic affinity of any unknown isolate can in principle be determined within the context of all previously sequenced 16S rRNAs without prior knowledge of what the organism is. In contrast to the original RNase T1 cataloging method, when digestion products are analyzed by mass spectrometry, products with the same base composition cannot be distinguished. Hence, it is possible that organisms that are not closely related (having different underlying sequences) might be falsely identified by mass spectral coincidence. We present a convenient spectral coincidence function for expressing the degree of similarity (or distance) between any two mass-spectra. Trees constructed using this function are consistent with those produced by direct comparison of primary sequences, demonstrating that the inherent degeneracy in mass spectrometric analysis of RNA fragments does not preclude correct organism identification. RESULTS: Neighbor-joining trees for important bacterial pathogens were generated using distances based on mass spectrometric observables and the spectral coincidence function. These trees demonstrate that most pathogens will be readily distinguished using mass spectrometric analyses of RNA digestion products. A more detailed, genus-level analysis of pathogens and near relatives was also performed, and it was found that assignments of genetic affinity were consistent with those obtained by direct sequence comparisons. Finally, typical values of the coincidence between organisms were also examined with regard to phylogenetic level and sequence variability. CONCLUSION: Cluster analysis based on comparison of mass spectrometric observables using the spectral coincidence function is an extremely useful tool for determining the genetic affinity of an unknown bacterium. Additionally, fragmentation patterns can determine within hours if an unknown isolate is potentially a known pathogen among thousands of possible organisms, and if so, which one.

Bacteria↗

Toward a catalog for the transcripts and proteins (sialome) from the salivary gland of the malaria vector Anopheles gambiae.

Hundreds of Anopheles gambiae salivary gland cDNA library clones have been sequenced. A cluster analysis based on sequence similarity at e(-60) grouped the 691 sequences into 251 different clusters that code for proteins with putative secretory, housekeeping, or unknown functions. Among the housekeeping cDNAs, we found sequences predicted to code for novel thioredoxin, tetraspanin, hemopexin, heat shock protein, and TRIO and MBF proteins. Among secreted cDNAs, we found 21 novel A. gambiae salivary sequences including those predicted to encode amylase, calreticulin, selenoprotein, mucin-like protein and 30-kDa allergen, in addition to antigen 5- and D7-related proteins, three novel salivary gland (SG)-like proteins and eight unique putative secreted proteins (Hypothetical Proteins, HP). The electronic version of this paper contains hyperlinks to FASTA-formatted files for each cluster with the best match to the nonredundant (NR) and conserved domain databases (CDD) in addition to CLUSTAL alignments of each cluster. The N terminus of 12 proteins (SG-1, SG-1-like 2, SG-6, HP 8, HP 9-like, 5' nucleotidase, 30-kDa protein, antigen 5- and four D7-related proteins) has been identified by Edman degradation of PVDF-transferred, SDS/PAGE-separated salivary gland proteins. Therefore, we contribute to the generation of a catalog of A. gambiae salivary transcripts and proteins. These data are freely available and will eventually become an invaluable tool to study the role of salivary molecules in parasite-host/vector interactions.

Amino Acid Sequence↗

Transcript annotation in FANTOM3: mouse gene catalog based on physical cDNAs.

The international FANTOM consortium aims to produce a comprehensive picture of the mammalian transcriptome, based upon an extensive cDNA collection and functional annotation of full-length enriched cDNAs. The previous dataset, FANTOM2, comprised 60,770 full-length enriched cDNAs. Functional annotation revealed that this cDNA dataset contained only about half of the estimated number of mouse protein-coding genes, indicating that a number of cDNAs still remained to be collected and identified. To pursue the complete gene catalog that covers all predicted mouse genes, cloning and sequencing of full-length enriched cDNAs has been continued since FANTOM2. In FANTOM3, 42,031 newly isolated cDNAs were subjected to functional annotation, and the annotation of 4,347 FANTOM2 cDNAs was updated. To accomplish accurate functional annotation, we improved our automated annotation pipeline by introducing new coding sequence prediction programs and developed a Web-based annotation interface for simplifying the annotation procedures to reduce manual annotation errors. Automated coding sequence and function prediction was followed with manual curation and review by expert curators. A total of 102,801 full-length enriched mouse cDNAs were annotated. Out of 102,801 transcripts, 56,722 were functionally annotated as protein coding (including partial or truncated transcripts), providing to our knowledge the greatest current coverage of the mouse proteome by full-length cDNAs. The total number of distinct non-protein-coding transcripts increased to 34,030. The FANTOM3 annotation system, consisting of automated computational prediction, manual curation, and final expert curation, facilitated the comprehensive characterization of the mouse transcriptome, and could be applied to the transcriptomes of other species.

Animals↗

[Localization of genes, determining quantitative traits in wheat: amendment to the "catalog of chromosomal mapping of genes in domestic cultivars of wheat"].

An amendment to the catalog of chromosome location of genes in Russian wheat cultivars was constructed with the published data of the recent decade. The results of chromosomal localization were summarized and analyzed by methods of multivariate statistics. Chromosomes critical for 40 quantitative traits under study proved to cluster according to their homeology, i.e., by homeological groups. The hypotheses providing an explanation for this finding are considered. It is suggested that quantitative traits are similarly controlled by genes located on homeological chromosomes in common wheat, making it possible to isolate a limited number of major genes for each particular quantitative trait.

Chromosome Mapping↗

A dedicated database program for cataloging recombinant clones and other laboratory products of molecular biology technology.

A novel computer database program dedicated to storing, cataloging, and accessing information about recombinant clones and libraries has been developed for the IBM (or compatible) personal computer. This program, named CLONES, also stores information about bacterial strains and plasmid and bacteriophage vectors used in molecular biology. The advantages of this method are improved organization of data, fast and easy assimilation of new data, automatic association of new data with existing data, and rapid retrieval of desired records using search criteria specified by the user. Individual records are indexed in the database using B-trees, which automatically index new entries and expedite later access. The use of multiple windows, pull-down menus, scrolling pick-lists, and field-input techniques make the program intuitive to understand and easy to use. Daughter databases can be created to include all records of a particular type, or only those records matching user-specified search criteria. Separate databases can also be merged into a larger database. This computer program provides an easy-to-use and accurate means to organize, maintain, access, and share information about recombinant clones and other laboratory products of molecular biology technology.

Database Management Systems↗

Standardized and extended catalog of major proteins of the human kidney.

A new version of a two-dimensional electrophoretic catalog of proteins of the human kidney and of various morphological and functional structures of the human kidney is presented; it contains information about 179 polypeptides. The following proteins were identified: crystallin, albumin, mitochondrial superoxide dismutase, actin, fatty acid-binding protein, alpha-ATP-synthase, and transferrin. Some protein groups are specific for certain morphological structures.

Amino Acid Sequence↗

[Information sources for the physician on the internet. Library catalogs, scientific bookshops and journal articles].

The fast availability of current medical literature is of paramount importance for physicians. Especially methods of online-retrieval in specialised libraries and of journal articles as well as the possibility of lending are of significance. Also, it should be possible to purchase books in online-bookshops. Therefore, the Internet with it's constantly increasing offers plays a more and more important role in scientific research.

Catalogs as Topic↗

Catalog of 300 SNPs in 23 genes encoding G-protein coupled receptors.

We previously published a series of detailed maps of single nucleotide polymorphisms (SNPs) in the genomic regions of 209 gene loci encoding drug metabolizing enzymes, transporters, receptors, and other potential drug targets. In addition to the maps reported earlier, we provide here high-resolution SNP maps of 23 genes encoding G-protein coupled receptors in the Japanese population. A total of 300 SNPs were identified through screening of these loci; 83 in four adenosine receptor family genes, 45 in three adrenergic receptor family genes, 22 in three EDG receptor family genes, 29 in three melanocortin receptor family genes, 22 in two somatostatin receptor family genes, 21 in five anonymous G protein-coupled receptor family genes, and 78 in the others (AVPR1B, OXTR, and TNFRSF1A). We also discovered a total of 33 genetic variations of other types. Of the 300 SNPs, 132 (44%) appeared to be novel on the basis of comparisons with the dbSNP database of the National Center for Biotechnology Information (US) or with previous publications. The maps constructed in this study will serve as an additional resource for studies of complex genetic diseases and drug-response phenotypes to be mapped by linkage-disequilibrium association analyses.

Catalogs as Topic↗

Catalog of 605 single-nucleotide polymorphisms (SNPs) among 13 genes encoding human ATP-binding cassette transporters: ABCA4, ABCA7, ABCA8, ABCD1, ABCD3, ABCD4, ABCE1, ABCF1, ABCG1, ABCG2, ABCG4, ABCG5, and ABCG8.

Single-nucleotide polymorphisms (SNPs) at some gene loci are useful as markers of individual risk for adverse drug reactions or susceptibility to complex diseases. We have been focusing on identifying SNPs in and around genes encoding drug-metabolizing enzymes and transporters, and have constructed several high-density SNP maps of such regions. Here we report SNPs at additional loci, specifically 13 genes belonging to the superfamily of ATP-binding cassette transporters ( ABCA4, ABCA7, ABCA8, ABCD1, ABCD3, ABCD4, ABCE1, ABCF1, ABCG1, ABCG2, ABCG4, ABCG5, and ABCG8). Sequencing a total of 416 kb of genomic DNA from 48 Japanese volunteers identified 605 SNPs among these 13 loci: 14 in 5' flanking regions, 5 in 5' untranslated regions, 37 within coding elements, 529 in introns, 8 in 3' untranslated regions, and 12 in 3' flanking regions. By comparing our data with SNPs deposited in the dbSNP database of the National Center for Biotechnology Information (US) and with published reports, we determined that 491 (81%) of the SNPs reported here were novel. We also detected 107 genetic variations of other types among the loci examined (insertion-deletions or mono- di-, or trinucleotide polymorphisms). The high-density SNP maps we constructed on the basis of these data should provide useful information for investigating associations between genetic variations and common diseases or responsiveness to drug therapy.

ATP-Binding Cassette Transporters↗