Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

A complete protein pattern of cellulase and hemicellulase genes in the filamentous fungus Trichoderma reesei.

The complete protein pattern of cellulase and hemicellulase genes was studied through the Genome-wide analysis in Trichoderma reesei. The genome database revealed the presence of 39 ORFs encoding related proteins, including 32 enzymes with a catalysis domain related to cellulases and hemicellulases and 7 related proteins with a cellulose-binding module (CBM). Ten of these encoded yet undescribed enzymes, including six novel beta-glucosidases or xylosidases, two putative xylanases and two undescribed mannases. To better illustrate the relation of these 39 related proteins, four groups were created and analyzed by phylogenetic analysis: group A corresponding to xylanases, group B belonging to mannases and acting to degrade mannan; group C containing all known and putative cellulose-degrading proteins that have highly conserved CBMs; and group D containing beta-glucosidase and beta-xylosidase. Group D was the largest group, in which 8 beta-glucosidases appeared to be non-secreted proteins.

Cellulase↗

The neurobeachin gene spans the common fragile site FRA13A.

Common fragile sites are normal constituents of chromosomal structure prone to chromosomal breakage. In humans, the cytogenetic locations of more than 80 common fragile sites are known. The DNA at 11 of them has been defined and characterized at the molecular level. According to the Genome Database, the common fragile site FRA13A maps to chromosome band 13q13.2. Here, we identify the precise genomic position of FRA13A, and characterize the genetic complexity of the fragile DNA sequence. We show that FRA13A breaks are limited to a 650 kb region within the neurobeachin (NBEA) gene, which genomically spans approximately 730 kb. NBEA encodes a neuron-specific multidomain protein implicated in membrane trafficking that is predominantly expressed in the brain and during development.

Autistic Disorder↗

Alu and L1 retroelements are correlated with the tissue extent and peak rate of gene expression, respectively.

We exploited the serial analysis of gene expression (SAGE) libraries and human genome database in silico to correlate the breadth of expression (BOE; housekeeping versus tissue-specific genes) and peak rate of expression (PRE; high versus low expressed genes) with the density distribution of the retroelements. The BOE status is linearly associated with the density of the sense Alus along the 100 kb nucleotides region upstream of a gene, whereas the PRE status is inversely correlated with the density of antisense L1s within a gene and in the up- and downstream regions of the 0-10 kb nucleotides. The radial distance of intranuclear position, which is known to serve as the global domain for transcription regulation, is reciprocally correlated with the fractions of Alu (toward the nuclear center) and L1 (toward the nuclear edge) elements in each chromosome. We propose that the BOE and PRE statuses are related to the reciprocal distribution of Alu and L1 elements that formulate local and global expression domains.

Alu Elements↗

OriDB: a DNA replication origin database.

Replication of eukaryotic chromosomes initiates at multiple sites called replication origins. Replication origins are best understood in the budding yeast Saccharomyces cerevisiae, where several complementary studies have mapped their locations genome-wide. We have collated these datasets, taking account of the resolution of each study, to generate a single list of distinct origin sites. OriDB provides a web-based catalogue of these confirmed and predicted S.cerevisiae DNA replication origin sites. Each proposed or confirmed origin site appears as a record in OriDB, with each record comprising seven pages. These pages provide, in text and graphical formats, the following information: genomic location and chromosome context of the origin site; time of origin replication; DNA sequence of proposed or experimentally confirmed origin elements; free energy required to open the DNA duplex (stress-induced DNA duplex destabilization or SIDD); and phylogenetic conservation of sequence elements. In addition, OriDB encourages community submission of additional information for each origin site through a User Notes facility. Origin sites are linked to several external resources, including the Saccharomyces Genome Database (SGD) and relevant publications at PubMed. Finally, a Chromosome Viewer utility allows users to interactively generate graphical representations of DNA replication data genome-wide. OriDB is available at www.oridb.org.

Chromosomes, Fungal↗

KinG: a database of protein kinases in genomes.

The KinG database is a comprehensive collection of serine/threonine/tyrosine-specific kinases and their homologues identified in various completed genomes using sequence and profile search methods. The database hosted at http://hodgkin. mbu.iisc.ernet.in/ approximately king provides the amino acid sequences, functional domain assignments and classification of gene products containing protein kinase domains. A search tool enabling the retrieval of protein kinases with specified subfamily and domain combinations is one of the key features of the resource. Identification of a kinase catalytic domain in the user's query sequence is possible using another search tool. The occurrence and location of critical catalytic residues if the query has a catalytic kinase domain, recognition of non-kinase domains in the sequence and subfamily classification of the kinase in the query will help in deciphering the biological role of the kinase. This online compilation can also be used to compare the protein kinases of a given subfamily and domain combinations across various genomes. Another exclusive feature of the database is the collection of the Ser/Thr/Tyr protein kinases and similar sequences encoded in the genomes of archaea and bacteria.

Animals↗

A searchable database for proteomes of oral microorganisms.

An online database of proteomes for two-dimensional electrophoresis (2DE) gel data was constructed and it is now freely accessible through a web-based interface. Proteins from three oral bacteria, Streptococcus mutans UA159, Actinobacillus actinomycetemcomitans HK1651, and Porphyromonas gingivalis W83, whose genome databases are freely available, were separated by 2DE, and protein spots were analyzed by matrix-assisted laser desorption/ionization time-of-flight (MALDI-TOF) and identified. About 1000 spots from the gels of P. gingivalis W83 were extracted and analyzed by MALDI-TOF, and 330 proteins were identified. In addition, 160 of 240 spots of A. actinomycetemcomitans and 158 of 356 spots of S. mutans were identified. Information such as spot coordinates on the gels, protein names (predicted functions), molecular weights, isoelectroric points, and links to online databases, including Oral Pathogen Sequence Databases of the Los Alamos National Laboratory Bioscience Division (ORALGEN) and National Center for Biotechnology Information (NCBI) or The Institute Genomic Research (TIGR), were stored in tables accessible through the relational database management system MySQL on an Apache web server. To test for functionality of this database system, responses of S. mutans to environmental changes were analyzed using the database and 21 spots on the gel were identified as proteins whose expression had been increased or decreased by environmental pH change without in-gel trypsin digestion, protein extraction, or MALDI-TOF/TOF-MS (mass spectrometer) analysis. The identified proteins are agreement with those reported in previous papers on acid tolerance of S. mutans, demonstrating the usefulness of the system. This database is available at http://www.myamagu.dent.kyushu-u.ac.jp/~bioinformatics/index.html or http://www.bipos.mascat.nihon-u.ac.jp/index.html.

Acids↗

dictyBase, the model organism database for Dictyostelium discoideum.

dictyBase (http://dictybase.org) is the model organism database (MOD) for the social amoeba Dictyostelium discoideum. The unique biology and phylogenetic position of Dictyostelium offer a great opportunity to gain knowledge of processes not characterized in other organisms. The recent completion of the 34 MB genome sequence, together with the sizable scientific literature using Dictyostelium as a research organism, provided the necessary tools to create a well-annotated genome. dictyBase has leveraged software developed by the Saccharomyces Genome Database and the Generic Model Organism Database project. This has reduced the time required to develop a full-featured MOD and greatly facilitated our ability to focus on annotation and providing new functionality. We hope that manual curation of the Dictyostelium genome will facilitate the annotation of other genomes.

Animals↗

Contribution of comparative fish studies to general endocrinology: structure and function of some osmoregulatory hormones.

Fish endocrinologists are commonly motivated to pursue their research driven by their own interests in these aquatic animals. However, the data obtained in fish studies not only satisfy their own interests but often contribute more generally to the studies of other vertebrates, including mammals. The life of fishes is characterized by the aquatic habitat, which demands many physiological adjustments distinct from the terrestrial life. Among them, body fluid regulation is of particular importance as the body fluids are exposed to media of varying salinities only across the thin respiratory epithelia of the gills. Endocrine systems play pivotal roles in the homeostatic control of body fluid balance. Judging from the habitat-dependent control mechanisms, some osmoregulatory hormones of fish should have undergone functional and molecular evolution during the ecological transition to the terrestrial life. In fact, water-regulating hormones such as vasopressin are essential for survival on the land, whereas ion-regulating hormones such as natriuretic peptides, guanylins and adrenomedullins are diversified and exhibit more critical functions in aquatic species. In this short review, we introduce some examples illustrating how comparative fish studies contribute to general endocrinology by taking advantage of such differences between fishes and tetrapods. In a functional context, fish studies often afford a deeper understanding of the essential actions of a hormone across vertebrate taxa. Using the natriuretic peptide family as an example, we suggest that more functional studies on fishes will bring similar rewards of understanding. At the molecular level, recent establishment of genome databases in fishes and mammals brings clues to the evolutionary history of hormone molecules via a comparative genomic approach. Because of the functional and molecular diversification of ion-regulating hormones in fishes, this approach sometimes leads to the discovery of new hormones in tetrapods as exemplified by adrenomedullin 2.

Adrenomedullin↗

Molecular cloning and characterization of a novel human putative transmembrane protein homologous to mouse sideroflexin associated with sideroblastic anemia.

Sideroflexin1 (Sfxn1), the prototype of a novel family of evolutionarily conserved proteins present in eukaryotes, has been found mutated in mice with siderocytic anemia. It is speculated that this protein facilitates the transport of a component required for iron utilization into mitochondrial. During the large-scale sequencing analysis of a human fetal brain cDNA library, we isolated a cDNA encoding a novel sideroflexin protein (SFXN4), which showed 59% identity and 71% similarity to mouse sideroflexin4. According to the search of the human genome database, SFXN4 gene is mapped to chromosome 10q25-26 and spans more than 24.7kb of the genomic DNA. It is 1428 base pair in length and the putative protein contains 305 amino acids with a conserved predicted five-transmembrane-domains structure. RT-PCR result shows that the SFXN4 gene is expressed in many tissues.

Amino Acid Sequence↗

Structure and chromosomal distribution of human mitochondrial pseudogenes.

Nuclear mitochondrial pseudogenes (Numts) have been found in the genome of many eukaryote species, including humans. Using a BLAST approach, we found 1105 DNA sequences homologous to mitochondrial DNA (mtDNA) in the August 2001 Goldenpath human genome database. We assembled these sequences manually into 286 pseudogenes on the basis of single insertion events and constructed a chromosomal map of these Numts. Some pseudogenes appeared highly modified, containing inversions, deletions, duplications, and displaced sequences. In the case of four randomly selected Numts, we used PCR tests on cells lacking mtDNA to ensure that our technique was free from genome-sequencing artifacts. Furthermore, phylogenetic investigation suggested that one Numt, apparently inserted into the nuclear genome 25-30 million years ago, had been duplicated at least 10 times in various chromosomes during the course of evolution. Thus, these pseudogenes should be very useful in the study of ancient mtDNA and nuclear genome evolution.

Chromosome Mapping↗

Chicken LRH-1 gene is transcribed from multiple promoters in steroidogenic organs.

Liver receptor homolog-1 (LRH-1) is a homolog of FTZ-F1, a transcription factor of the fruit fly, and belongs to the orphan nuclear receptor family. LRH-1 is expressed in organs derived from the endoderm, including intestine, liver and exocrine pancreas and plays a predominant role in development, bile-acid homeostasis, and reverse cholesterol transport. Recent research has revealed that mammalian LRH-1 is also expressed in the steroidogenic organs and has suggested that LRH-1 shares a role in steroidogenesis with steroidogenic factor-1 (SF-1), which is a paralog of LRH-1. In this study, we determined transcription initiation sites of chicken LRH-1 and showed that LRH-1 is expressed as several splicing variants in chicken steroidogenic organs. From three steroidogenic organs, the adrenal glands, ovaries, and testes, several cDNA fragments including different lengths and sequences were amplified by 5'-RACE and these were mainly classified into five types. Using these sequences, chicken genomic database was searched and four types of first exons were identified in chromosome 8. However, the database sequence of these regions included several gaps. So we cloned gap regions by PCR cloning from chicken genomic DNA and found the other type of first exons in the gaps. Moreover, RT-PCR showed the expression of LRH-1 in chicken steroidogenic organs as many splicing variants. We concluded that the chicken LRH-1 gene is transcribed from at least five different transcription initiation sites and alternative splicing produces several types of mRNA in steroidogenic organs.

Adrenal Glands↗

BarleyBase--an expression profiling database for plant genomics.

BarleyBase (BB) (www.barleybase.org) is an online database for plant microarrays with integrated tools for data visualization and statistical analysis. BB houses raw and normalized expression data from the two publicly available Affymetrix genome arrays, Barley1 and Arabidopsis ATH1 with plans to include the new Affymetrix 61K wheat, maize, soybean and rice arrays, as they become available. BB contains a broad set of query and display options at all data levels, ranging from experiments to individual hybridizations to probe sets down to individual probes. Users can perform cross-experiment queries on probe sets based on observed expression profiles and/or based on known biological information. Probe set queries are integrated with visualization and analysis tools such as the R statistical toolbox, data filters and a large variety of plot types. Controlled vocabularies for gene and plant ontologies, as well as interconnecting links to physical or genetic map and other genomic data in PlantGDB, Gramene and GrainGenes, allow users to perform EST alignments and gene function prediction using Barley1 exemplar sequences, thus, enhancing cross-species comparison.

Arabidopsis↗

Proteomic analysis on the expression of outer membrane proteins of Vibrio alginolyticus at different sodium concentrations.

The ability of osmoregulation is crucial to marine pathogens that always face the change of osmotic pressure when they shift between natural marine water-bodies and hosts. Previous studies indicated that the expressional patterns of outer membrane proteins (OMPs) changed when Gram-negative bacteria were transferred in different environments. In the present study, proteomic methodologies were used to investigate the expressional pattern of OMPs of Vibrio alginolyticus, a universal marine pathogen, at different Na(+) concentrations. OmpW, OmpV, and Omp TolC were determined to be osmotic stress responsive proteins. Of the three proteins, importantly, OmpV and OmpW showed distinctly reverse changes to each other, indicating that the two proteins might be the two components varied with changed NaCl concentrations. In addition, our results suggest that closely related species of bacteria with available whole genomic databases should be applied after item microorganism species was used when proteins from a bacterium with unavailable whole genomic information were identified by PMF. Therefore, our results not only expand our knowledge on osmotic stress responsive proteins, but also provide valuable information for strategies on screening of these proteins.

Amino Acid Sequence↗

Structure and regulation of acetyl-CoA carboxylase genes of metazoa.

Acetyl-CoA carboxylase (ACC) plays a fundamental role in fatty acid metabolism. The reaction product, malonyl-CoA, is both an intermediate in the de novo synthesis of long-chain fatty acids and also a substrate for distinct fatty acyl-CoA elongation enzymes. In metazoans, which have evolved energy storage tissues to fuel locomotion and to survive periods of starvation, energy charge sensing at the level of the individual cell plays a role in fuel selection and metabolic orchestration between tissues. In mammals, and probably other metazoans, ACC forms a component of an energy sensor with malonyl-CoA, acting as a signal to reciprocally control the mitochondrial transport step of long-chain fatty acid oxidation through the inhibition of carnitine palmitoyltransferase I (CPT I). To reflect this pivotal role in cell function, ACC is subject to complex regulation. Higher metazoan evolution is associated with the duplication of an ancestral ACC gene, and with organismal complexity, there is an increasing diversity of transcripts from the ACC paraloges with the potential for the existence of several isozymes. This review focuses on the structure of ACC genes and the putative individual roles of their gene products in fatty acid metabolism, taking an evolutionary viewpoint provided by data in genome databases.

Acetyl-CoA Carboxylase↗

Natural genetic variation for improving crop quality.

The narrow genetic basis of many crops combined with restrictions on the commercial use of genetically modified plants, has led to a surge of interest in exploring natural biodiversity as a source of novel alleles to improve the productivity, adaptation, quality and nutritional value of crops. Genetic methodologies have been applied to natural variation to improve quality aspects that are associated with the chemical composition of agricultural products. A future challenge in this emerging field is to integrate metabolic, phenotypic and genomic databases to allow a wider view of the plant metabolome and the application of this knowledge within genomics-assisted breeding.

Biodiversity↗

Seven evolutionarily conserved human rhodopsin G protein-coupled receptors lacking close relatives.

We report seven new members of the superfamily of human G protein-coupled receptors (GPCRs) found by searches in the human genome databases, termed GPR100, GPR119, GPR120, GPR135, GPR136, GPR141, and GPR142. We also report 16 orthologues of these receptors in mouse, rat, fugu (pufferfish) and zebrafish. Phylogenetic analysis shows that these are additional members of the family of rhodopsin-type GPCRs. GPR100 shows similarity with the orphan receptor SALPR. Remarkably, the other receptors do not have any close relative among other known human rhodopsin-like GPCRs. Most of these orphan receptors are highly conserved through several vertebrate species and are present in single copies. Analysis of expressed sequence tag (EST) sequences indicated individual expression patterns, such as for GPR135, which was found in a wide variety of tissues including eye, brain, cervix, stomach and testis. Several ESTs for GPR141 were found in marrow and cancer cells, while the other receptors seem to have more restricted expression patterns.

Amino Acid Sequence↗

Mapping Gene Ontology to proteins based on protein-protein interaction data.

MOTIVATION: Gene Ontology (GO) consortium provides structural description of protein function that is used as a common language for gene annotation in many organisms. Large-scale techniques have generated many valuable protein-protein interaction datasets that are useful for the study of protein function. Combining both GO and protein-protein interaction data allows the prediction of function for unknown proteins. RESULT: We apply a Markov random field method to the prediction of yeast protein function based on multiple protein-protein interaction datasets. We assign function to unknown proteins with a probability representing the confidence of this prediction. The functions are based on three general categories of cellular component, molecular function and biological process defined in GO. The yeast proteins are defined in the Saccharomyces Genome Database (SGD). The protein-protein interaction datasets are obtained from the Munich Information Center for Protein Sequences (MIPS), including physical interactions and genetic interactions. The efficiency of our prediction is measured by applying the leave-one-out validation procedure to a functional path matching scheme, which compares the prediction with the GO description of a protein's function from the abstract level to the detailed level along the GO structure. For biological process, the leave-one-out validation procedure shows 52% precision and recall of our method, much better than that of the simple guilty-by-association methods.

Chromosome Mapping↗

Physical linkage of expressed sequence tags (ESTs) to polymorphic markers on the X chromosome.

Expressed sequence tags (ESTs) can in principle serve as specialized sequence tagged sites (STSs) to assemble a functional map of the human genome. The strategy of physically linking ESTs to the nearest genetic linkage markers should provide specific candidate genes for the X-linked diseases associated with these markers or loci. Therefore, 19 ESTs assigned to the X chromosome in the Genome Database (GDB) were analyzed. Eighteen were confirmed to be X-specific and were localized to regions of the X chromosome using a panel of somatic cell hybrids. Localization was then refined by positioning them on yeast artificial chromosome (YAC)-based maps. Seventeen ESTs identified cognate YACs by PCR screening and 12 of the ESTs have been assembled in YAC contigs containing polymorphic and other X chromosomal markers. Two of them also produced syntenically equivalent products in mouse. Thus localizing ESTs relative to polymorphic markers will help to assemble an integrated physical and transcriptional map of the chromosome and provide candidates for disease-gene searches.

Animals↗