Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,027 records · Page 57Linked to original sources

Cloning and initial characterization of the Arabidopsis thaliana endoplasmic reticulum oxidoreductins.

The oxidation and isomerization of disulfide bonds is necessary for the growth of all organisms. In yeast, the oxidative folding of secretory pathway proteins is catalyzed by protein disulfide isomerase (PDI), which requires Ero1p (endoplasmic reticulum oxidoreductin) for its own oxidation. In Homo sapiens, two homologues of Ero1p, Ero1-Lalpha and Ero1-Lbeta, have been cloned. Both Ero1-Lalpha and Ero1-Lbeta interact via disulfide bonds with PDI and support the oxidation of immunoglobulin light chains. However, the function of Ero proteins in plants has not yet been analyzed. In this article, we report the cloning of the two Ero1p homologues present in Arabidopsis thaliana, demonstrating that one of the cDNAs has a shorter terminal exon than predicted and differs from the annotated sequence found in the genome database. Sequence analysis of the Arabidopsis endoplasmic reticulum oxidoreductins (AEROs) reveals that both AERO1 and AERO2 are more closely related to each other than to either of the human Eros. Both in vitro translated AERO proteins are targeted to the endoplasmic reticulum and glycosylated. The ability to use a genetically tractable multicellular organism in combination with biochemical approaches should further our understanding of redox networks and Ero function in both plants and animals.

Arabidopsis↗

Identification of CXorf1, a novel intronless gene in Xq27.3, expressed in human hippocampus.

We have identified and characterized a novel human gene (Nomenclature Committee of the Genome Database GDB-assigned symbol CXorf1) that maps to the long arm of the X chromosome in Xq27 between loci DXS369 and DXS181, approximately 2.5 Mb centromeric to the FMR1 gene. The CXorf1 gene is conserved in primates, cow, and horse but not in mouse and rat. Northern blot analysis revealed two transcripts, present in the brain and in the G361 melanoma cell line. In situ hybridization experiments performed on sections of human hippocampus showed a clear, uneven localization of the CXorf1 mRNA in specific subfields of this brain area. In particular, CXorf1 was localized in the granular-cell layer of the dentate gyrus and in the CA2-CA3 subfields of Ammon's horn. CXorf1 is one of the first genes from this region to be characterized in detail and, on the basis of its chromosomal location and expression pattern, may have an important function in the brain.

Amino Acid Sequence↗

Preparation and characterization of a monoclonal antibody against the protein LIGHT.

LIGHT (which is homologous to lymphotoxins, shows inducible expression, and competes with HSV glycoprotein D for HVEM, a receptor expressed by T lymphocytes [Genome Database designation, TNFSF14]), a newly identified member of the TNF superfamily, is up-regulated upon activation of T-cells. LIGHT plays an important role in the T-cell-mediated tumor and graft-versus-host disease via LIGHT/HVEM/LT beta R signaling. To prepare specific monoclonal antibody (MAb) against murine LIGHT, a fragment containing the extracellular domain of LIGHT was inserted into prokaryotic expression vector pET-32a(+). The his-tagged fusion protein was expressed in BL21(DE3) in the form of inclusion bodies. The fusion protein was purified and refolded on-column using immobilized mental affinity chromatography. Rat MAb against murine LIGHT was obtained with hybridoma technique and specific ELISA screening. Western blotting and flow cytometry assays showed that MAb 4C11 had specific binding ability with LIGHT protein in eukaryotic cells. Lymphocyte proliferation assays indicated that this MAb could co-stimulate the proliferation of T-cells. Thus, this MAb may be the basis for detection of LIGHT protein in tissue or cell and be beneficial for the study of LIGHT/HVEM/LT beta R pathway.

Animals↗

Clustering of chi sequence in Escherichia coli genome.

An 8-mer DNA sequence called chi (5'-GCTGGTGG) is present on the Escherichia coli chromosome at a high frequency. It is responsible for both the attenuation of RecBCD exonuclease activity and the promotion of RecABCD-mediated homologous recombination. chi was first identified as a site that increased plaque size of bacteriophage lambda. lambda containing chi makes very small plaques on a recC* (recC1004) mutant because chi is poorly recognized by the RecBC*D mutant enzyme. We cloned E. coli chromosomal fragments in lambda that allowed lambda to form larger plaques on this recC* mutant as well as on the rec+ parent. One identified fragment contained a cluster of two copies of chi and several chi-like sequences with the same orientation. It increased recombination in the rec+ strain more than a fragment with one chi did. This fragment was within the rep gene, whose helicase product is known to be required for growth in the absence of functional RecBCD enzyme. The possibility that RecBCD enzyme might interact both with the rep gene and its product is discussed. Many of the other chi clusters identified in the E. coli genome database lie within genes for membrane proteins. The possible significance of these findings is discussed.

Bacteriophage lambda↗

Deletion mutants in COP9/signalosome subunits in fission yeast Schizosaccharomyces pombe display distinct phenotypes.

The COP9/signalosome complex is highly conserved in evolution and possesses significant structural similarity to the 19S regulatory lid complex of the proteasome. It also shares limited similarity to the translation initiation factor eIF3. The signalosome interacts with multiple cullins in mammalian cells. In the fission yeast Schizosaccharomyces pombe, the Csn1 subunit is required for the removal of covalently attached Nedd8 from Pcu1, one of three S. pombe cullins. It remains unclear whether this activity is required for all the functions ascribed to the signalosome. We previously identified Csn1 and Csn2 as signalosome subunits in S. pombe. csn1 and csn2 null mutants are DNA damage sensitive and exhibit slow DNA replication. Two further putative subunits, Csn4 and Csn5, were identified from the S. pombe genome database. Herein, we characterize null mutations of csn4 and csn5 and demonstrate that both genes are required for removal of Nedd8 from the S. pombe cullin Pcu1 and that their protein products associate with Csn1 and Csn2. However, neither csn4 nor csn5 null mutants share the csn1 and csn2 mutant phenotypes. Our data suggest that the subunits of the signalosome cannot be considered as a distinct functional unit and imply that different subunits of the signalosome mediate distinct functions.

COP9 Signalosome Complex↗

Golgi localization of Syne-1.

We have previously identified a Golgi-localized spectrin isoform by using an antibody to the beta-subunit of erythrocyte spectrin. In this study, we show that a screen of a lambdagt11 expression library resulted in the isolation of an approximately 5-kb partial cDNA from a Madin-Darby bovine kidney (MDBK) cell line, which encoded a polypeptide of 1697 amino acids with low, but detectable, sequence homology to spectrin (37%). A blast search revealed that this clone overlaps with the 5' end of a recently identified spectrin family member Syne-1B/Nesprin-1beta, an alternately transcribed gene with muscle-specific forms that bind acetylcholine receptor and associate with the nuclear envelope. By comparing the sequence of the MDBK clone with sequence data from the human genome database, we have determined that this cDNA represents a central portion of a very large gene ( approximately 500 kb), encoding an approximately 25-kb transcript that we refer to as Syne-1. Syne-1 encodes a large polypeptide (8406 amino acids) with multiple spectrin repeats and a region at its amino terminus with high homology to the actin binding domains of conventional spectrins. Golgi localization for this spectrin-like protein was demonstrated by expression of epitope-tagged fragments in MDBK and COS cells, identifying two distinct Golgi binding sites, and by immunofluorescence microscopy by using several different antibody preparations. One of the Golgi binding domains on Syne-1 acts as a dominant negative inhibitor that alters the structure of the Golgi complex, which collapses into a condensed structure near the centrosome in transfected epithelial cells. We conclude that the Syne-1 gene is expressed in a variety of forms that are multifunctional and are capable of functioning at both the Golgi and the nuclear envelope, perhaps linking the two organelles during muscle differentiation.

Amino Acid Sequence↗

FORESST: fold recognition from secondary structure predictions of proteins.

MOTIVATION: A method for recognizing the three-dimensional fold from the protein amino acid sequence based on a combination of hidden Markov models (HMMs) and secondary structure prediction was recently developed for proteins in the Mainly-Alpha structural class. Here, this methodology is extended to Mainly-Beta and Alpha-Beta class proteins. Compared to other fold recognition methods based on HMMs, this approach is novel in that only secondary structure information is used. Each HMM is trained from known secondary structure sequences of proteins having a similar fold. Secondary structure prediction is performed for the amino acid sequence of a query protein. The predicted fold of a query protein is the fold described by the model fitting the predicted sequence the best. RESULTS: After model cross-validation, the success rate on 44 test proteins covering the three structural classes was found to be 59%. On seven fold predictions performed prior to the publication of experimental structure, the success rate was 71%. In conclusion, this approach manages to capture important information about the fold of a protein embedded in the length and arrangement of the predicted helices, strands and coils along the polypeptide chain. When a more extensive library of HMMs representing the universe of known structural families is available (work in progress), the program will allow rapid screening of genomic databases and sequence annotation when fold similarity is not detectable from the amino acid sequence. AVAILABILITY: FORESST web server at http://absalpha.dcrt.nih.gov:8008/ for the library of HMMs of structural families used in this paper. FORESST web server at http://www.tigr.org/ for a more extensive library of HMMs (work in progress). CONTACT: valedf@tigr.org; munson@helix.nih.gov; garnier@helix.nih.gov

Computer Simulation↗

Phylogenetic web profiler.

SUMMARY: Phylogenetic Web Profiler (PWP) is a web-based service designed to perform phylogenetic profiling of proteins against genomes. The current version offers a selection of 63 completed genomes and available plasmids as annotated in the PEDANT genome database. Unlike currently available applications, this tool offers several choices of ortholog prediction parameters including E-value cutoff, percent length difference tolerance, and annotation similarity. Additional features include tight integration with the PEDANT database and tools to analyze properties of predicted proteins. PWP should prove very useful for the analysis of functional-linkage between proteins.

Amino Acid Sequence↗

Predicting phenotype from patterns of annotation.

MOTIVATION: Predicting the outcome of specific experiments (such as the growth of a particular mutant strain in a particular medium) has the potential to allow researchers to devote resources to experiments with higher expected numbers of 'hits'. RESULTS: We use decision trees to predict phenotypes associated with Saccharomyces cerevisiae genes on the basis of Gene Ontology (GO) functional annotations from the Saccharomyces Genome Database (SGD) and other phenotypic annotations from the Yeast Phenotype Catalog at the Munich Information Center for Protein Sequences (MIPS). We assess the methodology in three ways: (1) we use cross-validation on the phenotypic annotations listed in MIPS, and show ROC curves indicating the tradeoff between true-positive rate and false-positive rate; (2) we do a literature-search for 100 of the predicted gene-phenotype associations that are not listed in MIPS, and find evidence for 43 of them; (3) we use deletion strains to experimentally assess 61 predicted gene-phenotype associations not listed in MIPS; significantly more of these deletion strains show abnormal growth than would be expected by chance.

Algorithms↗

Online synonymous codon usage analyses with the ade4 and seqinR packages.

UNLABELLED: Correspondence analysis of codon usage data is a widely used method in sequence analysis, but the variability in amino acid composition between proteins is a confounding factor when one wants to analyse synonymous codon usage variability. A simple and natural way to cope with this problem is to use within-group correspondence analysis. There is, however, no user-friendly implementation of this method available for genomic studies. Our motivation was to provide to the community a Web facility to easily study synonymous codon usage on a subset of data available in public genomic databases. AVAILABILITY: Availability through the Pole Bioinformatique Lyonnais (PBIL) Web server at http://pbil.univ-lyon1.fr/datasets/charif04/ with a demo allowing us to reproduce the figure in the present application note. All underlying software is distributed under a GPL licence. CONTACT: http://pbil.univ-lyon1.fr/members/lobry.

Algorithms↗

Detecting clusters of different geometrical shapes in microarray gene expression data.

MOTIVATION: Clustering has been used as a popular technique for finding groups of genes that show similar expression patterns under multiple experimental conditions. Many clustering methods have been proposed for clustering gene-expression data, including the hierarchical clustering, k-means clustering and self-organizing map (SOM). However, the conventional methods are limited to identify different shapes of clusters because they use a fixed distance norm when calculating the distance between genes. The fixed distance norm imposes a fixed geometrical shape on the clusters regardless of the actual data distribution. Thus, different distance norms are required for handling the different shapes of clusters. RESULTS: We present the Gustafson-Kessel (GK) clustering method for microarray gene-expression data. To detect clusters of different shapes in a dataset, we use an adaptive distance norm that is calculated by a fuzzy covariance matrix (F) of each cluster in which the eigenstructure of F is used as an indicator of the shape of the cluster. Moreover, the GK method is less prone to falling into local minima than the k-means and SOM because it makes decisions through the use of membership degrees of a gene to clusters. The algorithmic procedure is accomplished by the alternating optimization technique, which iteratively improves a sequence of sets of clusters until no further improvement is possible. To test the performance of the GK method, we applied the GK method and well-known conventional methods to three recently published yeast datasets, and compared the performance of each method using the Saccharomyces Genome Database annotations. The clustering results of the GK method are more significantly relevant to the biological annotations than those of the other methods, demonstrating its effectiveness and potential for clustering gene-expression data. AVAILABILITY: The software was developed using Java language, and can be executed on the platforms that JVM (Java Virtual Machine) is running. It is available from the authors upon request. SUPPLEMENTARY INFORMATION: Supplementary data are available at http://dragon.kaist.ac.kr/gk.

Algorithms↗

Protein classification using probabilistic chain graphs and the Gene Ontology structure.

MOTIVATION: Probabilistic graphical models have been developed in the past for the task of protein classification. In many cases, classifications obtained from the Gene Ontology have been used to validate these models. In this work we directly incorporate the structure of the Gene Ontology into the graphical representation for protein classification. We present a method in which each protein is represented by a replicate of the Gene Ontology structure, effectively modeling each protein in its own 'annotation space'. Proteins are also connected to one another according to different measures of functional similarity, after which belief propagation is run to make predictions at all ontology terms. RESULTS: The proposed method was evaluated on a set of 4879 proteins from the Saccharomyces Genome Database whose interactions were also recorded in the GRID project. Results indicate that direct utilization of the Gene Ontology improves predictive ability, outperforming traditional models that do not take advantage of dependencies among functional terms. Average increase in accuracy (precision) of positive and negative term predictions of 27.8% (2.0%) over three different similarity measures and three subontologies was observed. AVAILABILITY: C/C++/Perl implementation is available from authors upon request.

Algorithms↗

Integrating image data into biomedical text categorization.

Categorization of biomedical articles is a central task for supporting various curation efforts. It can also form the basis for effective biomedical text mining. Automatic text classification in the biomedical domain is thus an active research area. Contests organized by the KDD Cup (2002) and the TREC Genomics track (since 2003) defined several annotation tasks that involved document classification, and provided training and test data sets. So far, these efforts focused on analyzing only the text content of documents. However, as was noted in the KDD'02 text mining contest-where figure-captions proved to be an invaluable feature for identifying documents of interest-images often provide curators with critical information. We examine the possibility of using information derived directly from image data, and of integrating it with text-based classification, for biomedical document categorization. We present a method for obtaining features from images and for using them-both alone and in combination with text-to perform the triage task introduced in the TREC Genomics track 2004. The task was to determine which documents are relevant to a given annotation task performed by the Mouse Genome Database curators. We show preliminary results, demonstrating that the method has a strong potential to enhance and complement traditional text-based categorization methods.

Artificial Intelligence↗

Towards clustering of incomplete microarray data without the use of imputation.

MOTIVATION: Clustering technique is used to find groups of genes that show similar expression patterns under multiple experimental conditions. Nonetheless, the results obtained by cluster analysis are influenced by the existence of missing values that commonly arise in microarray experiments. Because a clustering method requires a complete data matrix as an input, previous studies have estimated the missing values using an imputation method in the preprocessing step of clustering. However, a common limitation of these conventional approaches is that once the estimates of missing values are fixed in the preprocessing step, they are not changed during subsequent processes of clustering; badly estimated missing values obtained in data preprocessing are likely to deteriorate the quality and reliability of clustering results. Thus, a new clustering method is required for improving missing values during iterative clustering process. RESULTS: We present a method for Clustering Incomplete data using Alternating Optimization (CIAO) in which a prior imputation method is not required. To reduce the influence of imputation in preprocessing, we take an alternative optimization approach to find better estimates during iterative clustering process. This method improves the estimates of missing values by exploiting the cluster information such as cluster centroids and all available non-missing values in each iteration. To test the performance of the CIAO, we applied the CIAO and conventional imputation-based clustering methods, e.g. k-means based on KNNimpute, for clustering two yeast incomplete data sets, and compared the clustering result of each method using the Saccharomyces Genome Database annotations. The clustering results of the CIAO method are more significantly relevant to the biological gene annotations than those of other methods, indicating its effectiveness and potential for clustering incomplete gene expression data. AVAILABILITY: The software was developed using Java language, and can be executed on the platforms that JVM (Java Virtual Machine) is running. It is available from the authors upon request.

Algorithms↗

Analysis of a human fungiform papillae cDNA library and identification of taste-related genes.

Various genes related to early events in human gustation have recently been discovered, yet a thorough understanding of taste transduction is hampered by gaps in our knowledge of the signaling chain. As a first step toward gaining additional insight, the expression specificity of genes in human taste tissue needs to be determined. To this end, a fungiform papillae cDNA library has been generated and analyzed. For validation of the library, taste-related gene probes were used to detect known molecules. Subsequently, DNA sequence analysis was performed to identify further candidates. Of 987 clones sequenced, clustering results in 288 contigs. Comparison of these contigs with genomic databases reveals that 207 contigs (71.9%) match known genes, 16 (5.6%) match hypothetical genes, eight (2.8%) match repetitive sequences and 57 (19.8%) have no or low similarity to annotated genes. The results indicate that despite a high level of redundancy, this human fungiform cDNA library contains specific taste markers and is valuable for investigation of both known and novel taste-related genes.

Computational Biology↗

Digital Kennison: A bioinformatics pipeline for rapid mapping of sequences to the Drosophila melanogaster Y chromosome.

The Drosophila melanogaster Y chromosome is currently known to contain 13 single-copy protein-coding genes, six of which are essential for male fertility, as well as several non-coding genes and abundant repetitive DNA. Localization of Y-linked sequences has traditionally relied on labor-intensive crosses using Kennison's translocation strains, which map Y-linked loci by generating flies deficient for each of the six Y-chromosome fertility regions (ks-1, ks-2, kl-1, kl-2, kl-3, and kl-5). Here we present Digital Kennison, a computational pipeline that recasts this classical mapping strategy as a sequence-based analysis. The pipeline queries eight genomic databases derived from Kennison's strains using BLAST and read coverage, assigning sequences to fertility regions or the centromeric region with a calibrated confidence score. We benchmarked the method on 60 Y-linked sequences spanning all seven regions, including single-copy protein-coding genes, Mst77Y family members, non-coding RNAs, and the centromere. Digital Kennison achieved 97% precision while resolving challenging cases, including boundary-spanning genes (PRY and Ppr-Y), fragmented Mst77Y copies, and FDY, which has a closely related autosomal paralog. Beyond validating known localizations, the pipeline localized the unmapped gene CG41561 to the kl-1region and reassigned the transcript CR40629-RC from the kl-2 region to kl-5. It also localized 7 of 16 recently transferred Y-linked sequences, including 4 with high confidence. Applied to 904 R6 scaffolds, Digital Kennison assigned 75% to fertility regions, including five currently annotated as autosomal-pericentromeric. Digital Kennison reduces sequence localization from weeks of genetic crosses to minutes of computation while preserving the power of classical translocation mapping.

Drosophila melanogaster↗

Large scale identification of genes involved in cell surface biosynthesis and architecture in Saccharomyces cerevisiae.

The sequenced yeast genome offers a unique resource for the analysis of eukaryotic cell function and enables genome-wide screens for genes involved in cellular processes. We have identified genes involved in cell surface assembly by screening transposon-mutagenized cells for altered sensitivity to calcofluor white, followed by supplementary screens to further characterize mutant phenotypes. The mutated genes were directly retrieved from genomic DNA and then matched uniquely to a gene in the yeast genome database. Eighty-two genes with apparent perturbation of the cell surface were identified, with mutations in 65 of them displaying at least one further cell surface phenotype in addition to their modified sensitivity to calcofluor. Fifty of these genes were previously known, 17 encoded proteins whose function could be anticipated through sequence homology or previously recognized phenotypes and 15 genes had no previously known phenotype.

Cell Membrane↗

Molecular cloning and characterization of an alpha1,3 fucosyltransferase, CEFT-1, from Caenorhabditis elegans.

We report on the identification, molecular cloning, and characterization of an alpha1,3 fucosyltransferase (alpha1,3FT) expressed by the nematode, Caenorhabditis elegans . Although C. elegans glycoconjugates do not express the Lewis x antigen Galbeta1-->4[Fucalpha1-->3]GlcNAcbeta-->R, detergent extracts of adult C.elegans contain an alpha1,3FT that can fucosylate both nonsialylated and sialylated acceptor glycans to generate the Lexand sialyl Lexantigens, as well as the lacdiNAc-containing acceptor GalNAcbeta1-->4GlcNAcbeta1-->R to generate GalNAcbeta1-->4 [Fucalpha1-->3]GlcNAcbeta1-->R. A search of the C.elegans genome database revealed the existence of a gene with 20-23% overall identity to all five cloned human alpha1,3FTs. The putative cDNA for the C.elegans alpha1,3FT (CEFT-1) was amplified by PCR from a cDNA lambdaZAP library, cloned, and sequenced. COS7 cells transiently transfected with cDNA encoding CEFT-1 express the Lex, but not sLexantigen. The CEFT-1 in the transfected cell extracts can synthesize Lex, but not sialyl Lex, using exogenous acceptors. A second fucosyltransferase activity was detected in extracts of C. elegans that transfers Fuc in alpha1,2 linkage to Gal specifically on type-1 chains. The discovery of alpha-fucosyltransferases in C. elegans opens the possibility of using this well-characterized nematode as a model system for studying the role of fucosylated glycans in the development and survival of C.elegans and possibly other helminths.

Amino Acid Sequence↗