Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

The mouse gene expression database GXD

The gene expression database (GXD) is being developed to store and integrate expression information for mouse development. GXD addresses many issues that apply to gene expression databases in general, and its data structures and supporting software tools are generalized in design and thus readily adaptable to other life stages and species. Integration of GXD with the mouse genome database (MGD) and interconnections with other relevant databases will place the gene expression data into the larger biological and analytical context. Here, we describe the design and implementation of GXD and illustrate, in particular, the gene expression annotator, an electronic system for submitting expression data to the database.Copyright 1997 Academic Press Limited Copyright 1997Academic Press Limited

Journal Article↗

Molecular cloning of the BCL-6 gene, a transcriptional repressor for B-cell differentiation, in torafugu (Takifugu rubripes).

B-cell lymphoma-6 (BCL-6) is a transcriptional repressor that prevents the terminal differentiation of mature B-cells to plasma cells, and is essential for germinal center formation in the primary lymphoid organs of mammals. In this study, we identified the BCL-6 gene in torafugu (Takifugu rubripes) using the torafugu genome database, and analyzed the expression of BCL-6 mRNA in various tissues of torafugu, using RT-PCR. The BCL-6 gene consisted of eight exons and seven introns spanning a genome of ca. 3.3 kb. BCL-6 mRNA contained a 2112 bp open reading frame encoding 703 amino acids, with a predicted protein size of 78.8 kDa. The predicted torafugu BCL-6 primary structure contains two conserved specific motifs, the BTB/POZ domain at the N-terminus and the sixC2H2-type zinc finger motifs at the C-terminal region. The homology of torafugu BCL-6 to those of zebrafish (Danio rerio), Xenopus laevis, mouse (Mus musculus) and human (Homo sapiens) is 76, 59, 60 and 60%, respectively. RT-PCR analysis revealed that BCL-6 mRNA is highly expressed in pronephros, thymus, intestine, ovary, brain, nasal cavity and muscle. These results imply that torafugu BCL-6 is involved in regulation of B-cell differentiation in torafugu.

Amino Acid Motifs↗

PIR: a new resource for bioinformatics.

UNLABELLED: The Protein Information Resource (PIR) has greatly expanded its Web site and developed a set of interactive search and analysis tools to facilitate the analysis, annotation, and functional identification of proteins. New search engines have been implemented to combine sequence similarity search results with database annotation information. The new PIR search systems have proved very useful in providing enriched functional annotation of protein sequences, determining protein superfamily-domain relationships, and detecting annotation errors in genomic database archives. AVAILABILITY: http://pir.georgetown.edu/. CONTACT: mcgarvey@nbrf.georgetown.edu

Animals↗

GXD: a Gene Expression Database for the laboratory mouse: current status and recent enhancements. The Gene Expresison Database group.

The Gene Expression Database (GXD) is a community resource of gene expression information for the laboratory mouse. The database is designed as an open-ended system that can integrate different types of expression data. New expression data are made available on a daily basis. Thus, GXD provides increasingly complete information about what transcripts and proteins are produced by what genes; where, when and in what amounts these gene products are expressed; and how their expression varies in different mouse strains and mutants. GXD is integrated with the Mouse Genome Database (MGD). Continuously refined interconnections with sequence databases and with databases from other species place the gene expression information in the larger biological and analytical context. GXD is accessible through the Mouse Genome Informatics Web site at http://www.informatics.jax.org/ or directly at http://www.informatics.jax.org/menus/expression_menu.shtm l

Alternative Splicing↗

Genomics of the ccoNOQP-encoded cbb3 oxidase complex in bacteria.

Many bacteria adapt to microoxic conditions by synthesizing a particular cytochrome c oxidase (cbb3) complex with a high affinity for O2, encoded by the ccoNOQP operon. A survey of genome databases indicates that ccoNOQP sequences are widespread in all sub-branches of Proteobacteria but otherwise are found only in bacteria of the CFB group ( Cytophaga, Flexibacter, Bacteroides). Our analysis of available genome sequences suggests four major strategies of regulating ccoNOQP expression in response to O2. The most widespread strategy involves direct regulation by the O2-responsive protein Fnr. The second strategy involves an O2-insensitive paralogue of Fnr, FixK, whose expression is regulated by the O2-responding FixLJ two-component system. A third strategy of mixed regulation operates in bacteria carrying both fnr and fixLJ-fixKgenes. Another, not yet identified, strategy is likely to operate in the epsilon-Proteobacteria Helicobacter pylori and Campylobacter jejuni which lack fnr and fixLJ-fixK genes. The FixLJ strategy appears specific for the alpha-subclass of Proteobacteria but is not restricted to rhizobia in which it was originally discovered.

Bacteria↗

Characterization of CA XV, a new GPI-anchored form of carbonic anhydrase.

The main function of CAs (carbonic anhydrases) is to participate in the regulation of acid-base balance. Although 12 active isoenzymes of this family had already been described, analyses of genomic databases suggested that there still exists another isoenzyme, CA XV. Sequence analyses were performed to identify those species that are likely to have an active form of this enzyme. Eight species had genomic sequences encoding CA XV, in which all the amino acid residues critical for CA activity are present. However, based on the sequence data, it was apparent that CA XV has become a non-processed pseudogene in humans and chimpanzees. RT-PCR (reverse transcriptase PCR) confirmed that humans do not express CA XV. In contrast, RT-PCR and in situ hybridization performed in mice showed positive expression in the kidney, brain and testis. A prediction of the mouse CA XV structure was performed. Phylogenetic analysis showed that mouse CA XV is related to CA IV. Therefore both of these enzymes were expressed in COS-7 cells and studied in parallel experiments. The results showed that CA XV shares several properties with CA IV, i.e. it is a glycosylated glycosylphosphatidylinositol-anchored membrane protein, and it binds CA inhibitor. The catalytic activity of CA XV is low, and the correct formation of disulphide bridges is important for the activity. Both specific and non-specific chaperones increase the production of active enzyme. The results suggest that CA XV is the first member of the alpha-CA gene family that is expressed in several species, but not in humans and chimpanzees.

Amino Acid Sequence↗

The spvB gene-product of the Salmonella enterica virulence plasmid is a mono(ADP-ribosyl)transferase.

A number of well-known bacterial toxins ADP-ribosylate and thereby inactivate target proteins in their animal hosts. Recently, several vertebrate ecto-enzymes (ART1-ART7) with activities similar to bacterial toxins have also been cloned. We show here that PSIBLAST, a position-specific-iterative database search program, faithfully connects all known vertebrate ecto-mono(ADP-ribosyl)transferases (mADPRTs) with most of the known bacterial mADPRTs. Intriguingly, no matches were found in the available public genome sequences of archaeabacteria, the yeast Saccharomyces cerevisiae or the nematode Caenorhabditis elegans. Significant new matches detected by PSIBLAST from the public sequence data bases included only one open reading frame (ORF) of previously unknown function: the spvB gene contained in the virulence plasmids of Salmonella enterica. Structure predictions of SpvB indicated that it is composed of a C-terminal ADP-ribosyltransferase domain fused via a poly proline stretch to a N-domain resembling the N-domain of the secretory toxin TcaC from nematode-infecting enterobacteria. We produced the predicted catalytic domain of SpvB as a recombinant fusion protein and demonstrate that it, indeed, acts as an ADP-ribosyltransferase. Our findings underscore the power of the PSIBLAST program for the discovery of new family members in genome databases. Moreover, they open a new avenue of investigation regarding salmonella pathogenesis.

ADP Ribose Transferases↗

Genomic circuits and the integrative biology of cardiac diseases.

Human cardiac disease is the result of complex interactions between genetic susceptibility and environmental stress. The challenge is to identify modifiers of disease, and to design new therapeutic strategies to interrupt the underlying disease pathways. The availability of genomic databases for many species is uncovering networks of conserved cardiac-specific genes within given physiological pathways. A new classification of human cardiac diseases can be envisaged based on the disruption of integrated genomic circuits that control heart morphogenesis, myocyte survival, biomechanical stress responses, cardiac contractility and electrical conduction.

Animals↗

Conserved modularity and potential for alternate splicing in mouse and human Slit genes.

The vertebrate Slit gene family currently consists of three members; Slit1, Slit2 and Slit3. Each gene encodes a protein containing multiple epidermal growth factor and leucine rich repeat motifs, which are likely to have importance in cell-cell interactions. In this study, we sought to fully define and characterise the vertebrate Slit gene family. Using long distance PCR coupled with in silico mapping, we determined the genomic structure of all three Slit genes in mouse and man. Analysis of EST and genomic databases revealed no evidence of further Slit family members in either organism. All three Slit genes were encoded by 36 (Slit3) or 37 (Slit1 and Slit2) exons covering at least 143 kb or 183 kb of mouse or human genomic DNA respectively. Two additional potential leucine-rich repeat encoding exons were identified within intron 12 of Slit2. These could be inserted in frame, suggesting that alternate splicing may occur in Slit2. A search for STS sequences within human Slit3 anchored this gene to D5S2075 at the 5' end (exon 4) and SGC32449 within the 3' UTR, suggesting that Slit3 may cover greater than 693 kb. The genomic structure of all Slit genes demonstrated considerable modularity in the placement of exon-intron boundaries such that individual leucine-rich repeat motifs were encoded by individual 72 bp exons. This further implies the potential generation of multiple Slit protein isoforms varying in their number of repeat units. cDNA library screening and EST database searching verified that such alternate splicing does occur.

3' Untranslated Regions↗

Reference Sequence Browser: An R application with a user-friendly GUI to rapidly query sequence databases.

Land managers, researchers, and regulators increasingly utilize environmental DNA (eDNA) techniques to monitor species richness, presence, and absence. In order to properly develop a biological assay for eDNA metabarcoding or quantitative PCR, scientists must be able to find not only reference sequences (previously identified sequences in a genomics database) that match their target taxa but also reference sequences that match non-target taxa. Determining which taxa have publicly available sequences in a time-efficient and accurate manner currently requires computational skills to search, manipulate, and parse multiple unconnected DNA sequence databases. Our team iteratively designed a Graphic User Interface (GUI) Shiny application called the Reference Sequence Browser (RSB) that provides users efficient and intuitive access to multiple genetic databases regardless of computer programming expertise. The application returns the number of publicly accessible barcode markers per organism in the NCBI Nucleotide, BOLD, or CALeDNA CRUX Metabarcoding Reference Databases. Depending on the database, we offer various search filters such as min and max sequence length or country of origin. Users can then download the FASTA/GenBank files from the RSB web tool, view statistics about the data, and explore results to determine details about the availability or absence of reference sequences.

User-Computer Interface↗

Sensitive protein comparisons with profiles and hidden Markov models.

Sequence database searches have become an important tool for the life sciences in general and for gene discovery-driven biotechnology in particular. Both the functional assignment of newly found proteins and the mining of genome databases for functional candidates are equally important tasks typically addressed by database searches. Sensitivity and reliability of the search methods are of crucial importance. The overall performance of sequence alignments and database searches can be enhanced considerably, when profiles or hidden Markov models (HMMs) derived from protein families are used as query objects instead of single sequences. This review discusses the concept of profiles, generalised profiles and profile-HMMs, the methods how they are constructed and the scope of possible applications in gene discovery and gene functional assignment.

Amino Acid Sequence↗

Integrating genomic knowledge sources through an anatomy ontology.

Modern genomic research has access to a plethora of knowledge sources. Often, it is imperative that researchers combine and integrate knowledge from multiple perspectives. Although some technology exists for connecting data and knowledge bases, these methods are only just beginning to be successfully applied to research in modem cell biology. In this paper, we argue that one way to integrate multiple knowledge sources is through anatomy--both generic cellular anatomy, as well as anatomic knowledge about the tissues and organs that may be studied via microarray gene expression experiments. We present two examples where we have combined a large ontology of human anatomy (the FMA) with other genomic knowledge sources: the gene ontology (GO) and the mouse genomic databases (MGD) of the Jackson Labs. These two initial examples of knowledge integration provide a proof of concept that anatomy can act as a hub through which we can usefully combine a variety of genomic knowledge and data.

Anatomy↗

Molecular cloning, gene organization and expression of the human UDP-GalNAc:Neu5Acalpha2-3Galbeta-R beta1,4-N-acetylgalactosaminyltransferase responsible for the biosynthesis of the blood group Sda/Cad antigen: evidence for an unusual extended cytoplasmic domain.

The nucleotide sequence of the short and long transcripts of beta1,4- N -acetylgalactosaminyltransferase have been submitted to the DDBJ, EMBL, GenBank(R) and GSDB Nucleotide Sequence Databases under accession nos AJ517770 and AJ517771 respectively. The human Sd(a) antigen is formed through the addition of an N -acetylgalactosamine residue via a beta1,4-linkage to a sub-terminal galactose residue substituted with an alpha2,3-linked sialic acid residue. We have taken advantage of the previously cloned mouse cDNA sequence of the UDP-GalNAc:Neu5Acalpha2-3Galbeta-R beta1,4- N -acetylgalactosaminyltransferase (Sd(a) beta1,4GalNAc transferase) to screen the human EST and genomic databases and to identify the corresponding human gene. The sequence spans over 35 kb of genomic DNA on chromosome 17 and comprises at least 12 exons. As judged by reverse transcription PCR, the human gene is expressed widely since it is detected in various amounts in almost all cell types studied. Northern blot analysis indicated that five Sd(a) beta1,4GalNAc transferase transcripts of 8.8, 6.1, 4.7, 3.8 and 1.65 kb were highly expressed in colon and to a lesser extent in kidney, stomach, ileum and rectum. The complete coding nucleotide sequence was amplified from Caco-2 cells. Interestingly, the alternative use of two first exons, named E1(S) and E1(L), leads to the production of two transcripts. These nucleotide sequences give rise potentially to two proteins of 506 and 566 amino acid residues, identical in their sequence with the exception of their cytoplasmic tail. The short form is highly similar (74% identity) to the mouse enzyme whereas the long form shows an unusual long cytoplasmic tail of 66 amino acid residues that is as yet not described for any other mammalian glycosyltransferase. Upon transient transfection in Cos-7 cells of the common catalytic domain, a soluble form of the protein was obtained, which catalysed the transfer of GalNAc residues to alpha2,3-sialylated acceptor substrates, to form the GalNAcbeta1-4[Neu5Acalpha2-3]Galbeta1-R trisaccharide common to both Sd(a) and Cad antigens.

Amino Acid Sequence↗

The Mouse Gene Expression Database (GXD).

The Gene Expression Database (GXD) is a community resource of gene expression information for the laboratory mouse. By combining the different types of expression data, GXD aims to provide increasingly complete information about the expression profiles of genes in different mouse strains and mutants, thus enabling valuable insights into the molecular networks that underlie normal development and disease. GXD is integrated with the Mouse Genome Database (MGD). Extensive interconnections with sequence databases and with databases from other species, and the development and use of shared controlled vocabularies extend GXD's utility for the analysis of gene expression information. GXD is accessible through the Mouse Genome Informatics web site at http://www.informatics.jax.org/ or directly at http://www.informatics.jax.org/menus/expression_menu. shtml.

Animals↗

Identification of the gene encoding the sole physiological fumarate reductase in Shewanella oneidensis MR-1.

Shewanella oneidensis MR-1 is a Gram-negative, nonfermentative rod with a complex electron transport system which facilitates its ability to use a variety of terminal electron acceptors, including fumarate, for anaerobic respiration. CMTn-3, a mutant isolated by transposon (TnphoA) mutagenesis, can no longer use fumarate as an electron acceptor; it lacks fumarate reductase activity as well as a 65-kDa soluble tetraheme flavocytochrome c. The sequence of the TnphoA-flanking genomic DNA of CMTn-3 did not align to those for fumarate reductase or related electron transport genes from other bacteria. Sequence analysis of the MR-1 genomic database demonstrated that an open reading frame encoding a 65-kDa tetraheme cytochrome c with sequence similarity to the fumarate reductase from S. frigidimarina NCIMB400 was found 8 kb away from the TnphoA-flanking genomic DNA of CMTn-3. PCR analysis demonstrated that a large deletion (>or=9.2 kb and <or=11 kb) of genomic DNA occurred in CMTn-3 as a result of TnphoA insertion. This deletion included at least half of the fumarate reductase gene as well as approximately 8 kb of upstream DNA. Complementation of CMTn-3 with the fumarate reductase gene plus 0.5-kb of upstream DNA restored growth on fumarate. These studies explicitly define the sole physiological fumarate reductase gene from the several possibilities suggested by the genomic sequence of MR-1. Surprisingly, the fumarate reductase gene plus 0.77-kb upstream DNA from S. frigidimarina NCIMB400 did not complement CMTn-3.

DNA Transposable Elements↗

Clustering time-varying gene expression profiles using scale-space signals.

The functional state of an organism is determined largely by the pattern of expression of its genes. The analysis of gene expression data from gene chips has primarily revolved around clustering and classification of the data using machine learning techniques based on the intensity of expression alone with the time-varying pattern mostly ignored. In this paper, we present a pattern recognition-based approach to capturing similarity by finding salient changes in the time-varying expression patterns of genes. Such changes can give clues about important events, such as gene regulation by cell-cycle phases, or even signal the onset of a disease. Specifically, we observe that dissimilarity between time series is revealed by the sharp twists and bends produced in a higher-dimensional curve formed from the constituent signals. Scale-space analysis is used to detect the sharp twists and turns and their relative strength with respect to the component signals is estimated to form a shape similarity measure between time profiles. A clustering algorithm is presented to cluster gene profiles using the scale-space distance as a similarity metric. Multi-dimensional curves formed from time series within clusters are used as cluster prototypes or indexes to the gene expression database, and are used to retrieve the functionally similar genes to a query gene profile. Extensive comparison of clustering using scale-space distance in comparison to traditional Euclidean distance is presented on the yeast genome database.

Algorithms↗

A web-based research tool for functional genomics of the microcirculation: the leukocyte adhesion cascade.

OBJECTIVE: Currently, microvascular data are almost exclusively deposited in scholarly journals. With the increasing availability of molecular data and the construction of genomic databases, a need arises to organize physiological data in a more accessible format. The microcirculation is a functional system that spans all organ systems and shares certain characteristics of organization with metabolic pathways, which have successfully been organized around genomic information for several prokaryotes. Here, we present a web-based research and teaching tool (http:@hsc.Virginia.EDU/medicine/basic-sci/biomed/l ey/) that covers a small aspect of microcirculatory physiology, the leukocyte adhesion cascade. METHODS: Currently, the site is organized in a flat-text mode with hypertext links to GenBank, Medline, and online journals where available. RESULTS: The web-based research and education tool is useful for graduate student education, and as a research resource for genetic researchers interested in gene function and for physiologists and biomedical engineers interested in the molecular basis of the leukocyte adhesion cascade. Our effort is intended to be a beginning toward a distributed database for the microcirculation (microcirculation physiome. see http://www.bme.jhu.edu/news/microphys). CONCLUSIONS: We invite all microvascular researchers to provide annotations and comments to make the research and teaching tool more useful to the scientific community.

Algorithms↗

Improving gene annotation using peptide mass spectrometry.

Annotation of protein-coding genes is a key goal of genome sequencing projects. In spite of tremendous recent advances in computational gene finding, comprehensive annotation remains a challenge. Peptide mass spectrometry is a powerful tool for researching the dynamic proteome and suggests an attractive approach to discover and validate protein-coding genes. We present algorithms to construct and efficiently search spectra against a genomic database, with no prior knowledge of encoded proteins. By searching a corpus of 18.5 million tandem mass spectra (MS/MS) from human proteomic samples, we validate 39,000 exons and 11,000 introns at the level of translation. We present translation-level evidence for novel or extended exons in 16 genes, confirm translation of 224 hypothetical proteins, and discover or confirm over 40 alternative splicing events. Polymorphisms are efficiently encoded in our database, allowing us to observe variant alleles for 308 coding SNPs. Finally, we demonstrate the use of mass spectrometry to improve automated gene prediction, adding 800 correct exons to our predictions using a simple rescoring strategy. Our results demonstrate that proteomic profiling should play a role in any genome sequencing project.

Algorithms↗