Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “bioinformatic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Genome data mining of lactic acid bacteria: the impact of bioinformatics.

Lactic acid bacteria (LAB) have been widely used in food fermentations and, more recently, as probiotics in health-promoting food products. Genome sequencing and functional genomics studies of a variety of LAB are now rapidly providing insights into their diversity and evolution and revealing the molecular basis for important traits such as flavor formation, sugar metabolism, stress response, adaptation and interactions. Bioinformatics plays a key role in handling, integrating and analyzing the flood of 'omics' data being generated. Reconstruction of metabolic potential using bioinformatics tools and databases, followed by targeted experimental verification and exploration of the metabolic and regulatory network properties, are the present challenges that should lead to improved exploitation of these versatile food bacteria.

Adaptation, Biological↗

Mitochondrial DNA identification of game and harvested freshwater fish species.

The use of DNA in forensics has grown rapidly for human applications along with the concomitant development of bioinformatics and demographic databases to help fully realize the potential of this molecular information. Similar techniques are also used routinely in many wildlife cases, such as species identification in food products, poaching and the illegal trade of endangered species. The use of molecular techniques in forensic cases related to wildlife and the development of associated databases has, however, mainly focused on large mammals with the exception of a few high-profile species. There is a need to develop similar databases for aquatic species for fisheries enforcement, given the large number of exploited and endangered fish species, the intensity of exploitation, and challenges in identifying species and their derived products. We sequenced a 500bp fragment of the mitochondrial cytochrome b gene from representative individuals from 26 harvested fish taxa from Ontario, Canada, focusing on species that support major commercial and recreational fisheries. Ontario provides a unique model system for the development of a fish species database, as the province contains an evolutionarily diverse array of freshwater fish families representing more than one third of all freshwater fish in Canada. Inter- and intraspecific sequence comparisons using phylogenetic analysis and a BLAST search algorithm provided rigorous statistical metrics for species identification. This methodology and these data will aid in fisheries enforcement, providing a tool to easily and accurately identify fish species in enforcement investigations that would have otherwise been difficult or impossible to pursue.

Animals↗

Characterization of isoforms and genomic organization of mouse calumenin.

Calumenin is a multiple EF-hand protein located in endo/sarcoplasmic reticulum of mammalian heart and other tissues [J. Biol. Chem. 272 (1997) 18232; Genomics 49 (1998) 331; Biochim. Biophys. Acta 1386 (1998) 121]. In the present study, a new isoform of mouse calumenin (mouse calumenin 2) was cloned by RT-PCR and genomic DNA PCR. The deduced amino acid sequence of mouse calumenin 2 is 315 aa long with the calculated MW of 37,064 and pI of 4.26. It has 92% aa sequence identity to previously identified mouse calumenin [J. Biol. Chem. 272 (1997) 18232] (mouse calumenin 1). The difference in the aa sequence was restricted to the first two EF-hand regions (residues 74-138). Northern blot analysis shows that mouse calumenin 2 is highly expressed in heart, lung, testis and unpregnant uterus. The expression of mouse calumenin 2 appears to decrease when fetal development is progressed. Genomic DNA PCR, sequencing and data mining of mouse genome database were utilized to examine the exon-intron boundaries of mouse calumenin genes. Both mouse calumenin 1 and 2 genes encompass six exons, and five of them (Exon1, 3, 4, 5 and 6) are identical. However, mouse calumenin 1 contains Exon2-1, whereas mouse calumenin 2 contains a neighboring Exon2-2. The calumenin genes are localized on mouse chromosome 6 having conserved synteny with human chromosome 7q32. For comparison, the genomic organization of human calumenin was also examined using the published human genome database (UCSC Genome Bioinformatics at ). Like mouse calumenin genes, two human calumenin genes also consist of five identical exons (Exon1, 3, 4, 5 and 6) and a different Exon2. The present study suggests that the genomic organization of calumenin genes is well conserved between human and mouse.

Amino Acid Sequence↗

Development of the polymerase chain reaction assay based on the canine genome database for detection of monoclonality in B cell lymphoma.

From the canine genome database and its bioinformatic analysis, we identified conserved sequences within the vast majority of 61 variable segments and 1 joining segment of the immunoglobulin heavy chain (IgH) gene, and designed optimal primers for polymerase chain reaction (PCR) amplification directed at these conserved sequences to evaluate the monoclonality of IgH in canine B cell lymphoma. Using the primers, a PCR-based assay was performed on fine-needle aspiration samples of normal, hyperplasia, and malignant lymph nodes and lymphoma cell lines. All fine-needle aspiration samples of five B cell lymphoma cases and the B cell lymphoma line GL-1 exhibited clonal amplification, whereas no amplification was observed in the samples from normal and hyperplasia lymph nodes, cases of T cell lymphoma, and the T cell lymphoma line CL-1. The primers we designed clearly distinguished malignant B lymphocytes from normal, reactive, and malignant T lymphocytes, indicating a potential utility of the primers for PCR-based routine clinical examination for canine B cell lymphoma.

Animals↗

Using molecular techniques to identify new microbial biocatalysts.

Evolution has favoured microorganisms that produce efficient enzymes with substrate-adapted biocatalytic activities. Progress in molecular techniques, especially expression cloning, molecular screening, protein engineering and in vivo and in vitro shuffling, have paved the way for greater speed and accuracy in cloning enzyme genes from microorganisms and generating versions with improved properties. Recently, two new approaches have been added: screening directly from uncultivated microorganisms and generating additional hits by database mining using bioinformatic tools.

Amino Acid Sequence↗

Construction of a human glycogene library and comprehensive functional analysis.

Eighteen years have passed after the first mammalian glycosyltransferase was cloned. At the beginning of April, 2001, 110 genes for human glycosyltransferases, including modifying enzymes for carbohydrate chains such as sulfotransferases, had been cloned and analyzed. We started the Glycogene Project (GG project) in April 2001, a comprehensive study on human glycogenes with the aid of bioinformatic technology. The term glycogene includes the genes for glycosyltransferases, sulfotransferases adding sulfate to carbohydrates and sugar-nucleotide transporters, etc. Firstly, as many novel genes, which are the candidates for glycogenes, as possible were searched using bioinformatic technology in databases. They were then cloned and expressed in various expression systems to detect the activity for carbohydrate synthesis. Their substrate specificity was determined using various acceptors.

Animals↗

Gene expression profiling in the human hypothalamus-pituitary-adrenal axis and full-length cDNA cloning.

The primary neuroendocrine interface, hypothalamus and pituitary, together with adrenals, constitute the major axis responsible for the maintenance of homeostasis and the response to the perturbations in the environment. The gene expression profiling in the human hypothalamus-pituitary-adrenal axis was catalogued by generating a large amount of expressed sequence tags (ESTs), followed by bioinformatics analysis (http://www.chgc.sh.cn/ database). Totally, 25,973 sequences of good quality were obtained from 31,130 clones (83.4%) from cDNA libraries of the hypothalamus, pituitary, and adrenal glands. After eliminating 5,347 sequences corresponding to repetitive elements and mtDNA, 20,626 ESTs could be assembled into 9, 175 clusters (3,979, 3,074, and 4,116 clusters in hypothalamus, pituitary, and adrenal glands, respectively) when overlapping ESTs were integrated. Of these clusters, 2,777 (30.3%) corresponded to known genes, 4,165 (44.8%) to dbESTs, and 2,233 (24.3%) to novel ESTs. The gene expression profiles reflected well the functional characteristics of the three levels in the hypothalamus-pituitary-adrenal axis, because most of the 20 genes with highest expression showed statistical difference in terms of tissue distribution, including a group of tissue-specific functional markers. Meanwhile, some findings were made with regard to the physiology of the axis, and 200 full-length cDNAs of novel genes were cloned and sequenced. All of these data may contribute to the understanding of the neuroendocrine regulation of human life.

Alternative Splicing↗

Bioinformatics in drug development and assessment.

Bioinformatics is playing an increasingly important role in nearly all aspects of drug discovery, drug assessment, and drug development. This growing importance lies not only in the role that bioinformatics plays in handling large volumes of data, but also in the utility of bioinformatics tools to predict, analyze, or help interpret clinical and preclinical findings. This review focuses on describing and evaluating some of the newer or more important bioinformatics resources (i.e., databases and software) that are of growing importance to understanding or predicting drug metabolism, especially with respect to the absorption, distribution, metabolism, excretion, (ADME), and toxicity (T) of both existing drugs and potential drug leads. Detailed descriptions and critical assessments of a number of potentially useful bioinformatics/cheminformatics databases and predictive ADMET software tools are provided. Additionally, several pharmaceutically important applications of both the databases and software are highlighted. Given the rapid growth in this area and the rapid changes that are taking place, a special emphasis is placed on freely available or Web-accessible resources.

Animals↗

Genome-wide identification and molecular characterization of Ole_e_I, Allerg_1 and Allerg_2 domain-containing pollen-allergen-like genes in Oryza sativa.

Pollen allergens play important roles in plant development in addition to their allergenic nature for human. More than 10 groups of pollen allergens have been reported. Among them, Pollen_Ole_e_I (Ole), Pollen_allerg_1 (Allerg1) and Pollen_allerg_2 (Allerg2) domain-containing proteins are the majority of allergens. We have identified 114 pollen-allergen-like genes in rice genome by bioinformatics using public databases. Among them, 45 genes encode Ole domain-containing proteins, 62 with Allerg1 and 7 with Allerg2. They are distributed on 11 of 12 rice chromosomes excluding chromosome 11. Comparison analysis of coding regions from both predicted genes and isolated full-length cDNAs showed that most of predicted genes were correct in the splicing of exons and introns, and only 7 exhibited wrong predictions. The fact suggested the applicability of the prediction programs to identify pollen-allergen genes. Phylogenetic analysis revealed the high diversity within OsOle genes and recent evolutionary event in OsAllerg1 genes, and suggested that some of OsOle genes were new members of the family. Expression analysis by RT-PCR showed that most of the genes were expressed in all tested tissues and only eight genes exhibited panicle-specific expression, suggesting that pollen-allergen genes play roles in not only productive but also vegetative development.

Allergens↗

High-throughput proteomics for alcohol research.

This report summarizes the proceedings of a satellite symposium of the 2003 Research Society on Alcoholism meeting held on June 21, 2003, in Fort Lauderdale, FL. The goal of this symposium, sponsored by the NIAAA, was to identify new proteomic directions in alcohol research that will (1) enable studies that focus on characterizing protein function, biochemical pathways, and networks to understand alcohol-related illnesses; (2) identify protein-protein interactions, posttranslational modifications, and subcellular localizations; (3) identify molecular targets for medication development; (4) develop biomarkers for susceptibility, dependence, consumption, and relapse, as well as alcohol-induced pathologies; and (5) develop high-throughput drug screens to test the efficacy of therapeutics that control alcohol-induced diseases. The purpose of the symposium was also to promote the application of high-throughput proteomic approaches, including isolation of membrane-bound proteins, in situ proteomics, large-scale two-dimensional separations, protein microarray platforms, mass spectrometry, matrix-assisted laser desorption/ionization, matrix-assisted laser desorption/ionization time-of-flight, liquid chromatography-tandem mass spectrometry, and isotope-coded affinity tags. In addition, the development of protein network maps by using new bioinformatics approaches for database mining was also discussed.

Alcohol Drinking↗

From ORFeome to biology: a functional genomics pipeline.

As several model genomes have been sequenced, the elucidation of protein function is the next challenge toward the understanding of biological processes in health and disease. We have generated a human ORFeome resource and established a functional genomics and proteomics analysis pipeline to address the major topics in the post-genome-sequencing era: the identification of human genes and splice forms, and the determination of protein localization, activity, and interaction. Combined with the understanding of when and where gene products are expressed in normal and diseased conditions, we create information that is essential for understanding the interplay of genes and proteins in the complex biological network. We have implemented bioinformatics tools and databases that are suitable to store, analyze, and integrate the different types of data from high-throughput experiments and to include further annotation that is based on external information. All information is presented in a Web database (http://www.dkfz.de/LIFEdb). It is exploited for the identification of disease-relevant genes and proteins for diagnosis and therapy.

Animals↗

The biology of chemokines and their receptors.

During the last five years, the development of bioinformatics and EST databases has been primarily responsible for the identification of many new chemokines and chemokine receptors. The chemokine field has also received considerable attention since chemokine receptors were found to act as co-receptors for HIV infection (1). In addition, chemokines, along with adhesion molecules, are crucial during inflammatory responses for a timely recruitment of specific leukocyte subpopulations to sites of tissue damage. However, chemokines and their receptors are also important in dendritic cell maturation (2), B (3), and T (4) cell development, Th1 and Th2 responses, infections, angiogenesis, and tumor growth as well as metastasis (5). Furthermore, an increase in the number of chemokine/receptor transgenic and knock-out mice has helped to define the functions of chemokines in vivo. In this review we discuss some of the chemokines' biological effects in vivo and in vitro, described in the last few years, and the implications of these findings when considering chemokine receptors as therapeutic targets.

Animals↗

Evaluation of common gene expression patterns in the rat nervous system.

In the postgenomic era, integrating data obtained from array technologies (e.g., oligonucleotide microarrays) with published information on eukaryotic genomes is beginning to yield biomarkers and therapeutic targets that are key for the diagnosis and treatment of disease. Nevertheless, identifying and validating these drug targets has not been a trivial task. Although a plethora of bioinformatics tools and databases are available, major bottlenecks for this approach reside in the interpretation of vast amounts of data, its integration into biologically representative models, and ultimately the identification of pathophysiologically and therapeutically useful information. In the field of neuroscience, accomplishing these goals has been particularly challenging because of the complex nature of nerve tissue, the relatively small adaptive nature of induced-gene expression changes, as well as the polygenic etiology of most neuropsychiatric diseases. This report combines published data sets from multiple transcript profiling studies that used GeneChip microarrays to illustrate a postanalysis approach for the interpretation of data from neuroscience microarray studies. By defining common gene expression patterns triggered by diverse events (administration of psychoactive drugs and trauma) in different nerve tissues (telencephalic brain areas and spinal cord), we broaden the conclusions derived from each of the original studies. In addition, the evaluation of the identified overlapping gene lists provides a foundation for generating hypotheses relating alterations in specific sets of genes to common physiological processes. Our approach demonstrates the significance of interpreting transcript profiling data within the context of common pathways and mechanisms rather than specific to a given tissue or stimulus. We also highlight the use of gene expression patterns in predictive biology (e.g., in toxicogenomics) as well as the utility of combining data derived from multiple microarray studies that examine diverse biological events for a broader interpretation of data from a particular microarray study.

Animals↗

Use of robust-long serial analysis of gene expression to identify novel fungal and plant genes involved in host-pathogen interactions.

Identification of important transcripts from fungal pathogens and host plants is indispensable for full understanding the molecular events occurring during fungal-plant interactions. Recently, we developed an improved LongSAGE method called robust-long serial analysis of gene expression (RL-SAGE) for deep transcriptome analysis of fungal and plant genomes. Using this method, we made 10 RL-SAGE libraries from two plant species (Oryza sativa and Zea maize) and one fungal pathogen (Magnaporthe grisea). Many of the transcripts identified from these libraries were novel in comparison with their corresponding EST collections. Bioinformatic tools and databases for analyzing the RL-SAGE data were developed. Our results demonstrate that RL-SAGE is an effective approach for large-scale identification of expressed genes in fungal and plant genomes.

Base Sequence↗

[Structural features of GR6 gene and its expression in colorectal neoplasm].

OBJECTIVE: To determine the characteristics of GR6 gene and putative GR6 protein, and to evaluate the expression of GR6 gene in colorectal cancer and normal mucosa. METHODS: Bioinformatic software and databases were applied to analyze the characteristics of GR6 gene and putative GR6 protein. Semi-quantitative RT-PCR was used to detect the expression of GR6 gene in colorectal carcinoma, adenoma and normal mucosa. RESULT: GR6 gene, encoding putative GR6 protein, consisted of 3 exons and contained 4 CpG islands by sequence analysis. It was predicted that putative GR6 protein included one protein kinase C phosphorylation site, one casein kinase II phosphorylation site, and three N-myristoylation sites. PSORT II software analysis predicted that putative GR6 protein was located in nucleus (reliability: 76.7%). At the level of mRNA, the expression of GR6 gene was high in normal mucosa, moderate in mucosa adjacent to cancer and adenoma tissue, low in colorectal carcinoma tissue. Significant differences were demonstrated between normal mucosa and adenoma (P<0.05), normal mucosa and carcinoma (P<0.01). CONCLUSION: The putative GR6 protein, encoded by GR6 gene may predictably function as an important nuclear signal transduction molecule. Decreased expression of GR6 gene may play an important role in the initiation and promotion of colorectal neoplasia.

Amino Acid Sequence↗

On the way to functional genomics.

Proteomics, the global comprehensive analysis of cellular proteins, will contribute to our understanding of gene function in the post-genomic era. The strategies of proteome analysis include the protein expression proteomics dealing with the comparison of cellular or tissue protein levels in the control and affected state (e.g. in the disease) and cell mapping proteomics aimed to define protein-protein interactions that constitute intracellular signalling network. The currently used methods of proteome analysis represent two-dimensional gel electrophoresis for protein separation, the image analysis means for protein maps comparison and the mass spectrometry (MALDI-TOF, MALDI-ESI) and bioinformatics means (sequence databases) for protein identification. For the protein interaction studies the yeast two-hybrid system is mostly employed. New concepts of disease diagnosis and treatment as well as new drugs design can be envisioned from the proteomic analysis.

Genomics↗

[Gene screening of dermal papilla cells in the state of aggregative growth and full-length cloning of HSPC016 gene for bioinformatics].

OBJECTIVE: To screen the genes of dermal papilla cells (DPC) related to the property of aggregative growth, and clone the full-length cDNA of differential HSPC016 gene for functional analysis. METHODS: DPC were collected from the hair of an individual aged 18 approximately 30 6 hours after the death. The complete papillae of hair were isolated and then cultured. The total DNA was extracted. Suppression subtractive hybridization-polymerase chain reaction was employed to screen out the genes differentially expressed in the DPC under the state of aggregative growth pattern in vitro. Then, rapid amplification of cDNA ends (RACE) technique was used to amplify the full-length cDNA of HSPC016 in the DPC, and bioinformatic methods were used to analyze their possible function. RESULTS: A subtractive library of human DPC was set up, and some up-regulated and down-regulated genes in the DPCs were screened out successfully. HSPC016 was identified to be a gene of 400 bp cDNA. Bioinformatic analyses and databases searching on the Internet indicated that this gene was mapped on chromosome 3 q21.31 and included an open reading frame with 195 bp coding for an expected 64aa soluble protein. The putative protein belonged to PD053992 family and was homologous to T2FA gene in domain. CONCLUSION: The establishment of subtractive library of DPC provides a solid foundation for further screening the genes related to aggregative growth and analyzing their regulatory mechanisms in DPC. Further more, HSPC016 in DPC may act as a subunit of a functional complex and play a role on transcriptional regulation within nucleus.

Adult↗

The Radiation Hybrid Database.

Since July 1995, the European Bioinformatics Institute (EBI) has maintained RHdb (http://www.ebi.ac.uk/RHdb/RHdb.html ), a public database for radiation hybrid data. Radiation hybrid mapping is an important technique for determining high resolution maps. Recently, CORBA access has been added to Rhdb. The EBI is an Outstation of the European Molecular Biology Laboratory (EMBL).

Animals↗