Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Screening and identification of key genes related to the immune microenvironment of rectal cancer influenced by radiotherapy based on bioinformatics methods.

OBJECTIVE: Radiotherapy (RT) plays a crucial role in the comprehensive treatment of rectal cancer. However, the impact of radiotherapy on the tumor microenvironment (TME), especially its effect on immune cell infiltration and immune-related gene expression, has not been fully studied. This study aims to screen and analyze key genes related to the immune microenvironment of rectal cancer influenced by radiotherapy based on bioinformatics methods for the purpose of identifying potential biomarkers and providing new insights for the personalized therapy of rectal cancer. METHODS: Using data from the Public Gene Expression Database (GEO) and the Cancer Genomics Database (TCGA), the impact of radiotherapy on the immune microenvironment of rectal cancer was explored using bioinformatics tools. Through screening differentially expressed genes (DEGs), correlation analysis, TIMER database analysis, immune infiltration score, and correlation analysis between key genes and prognosis, the effects of radiotherapy on the immune microenvironment of rectal cancer were investigated. RESULTS: Totally 7 upregulated and 4 downregulated differentially expressed genes were identified, among which MASP1, LTK, SLC9A3R2 were negatively correlated with myeloid suppressor cell infiltration (MDSCs), while ZP2 was positively correlated. The expression of MASP1 and SLC9A3R2 was closely related to the level of immune cell infiltration and played significant roles in the immune microenvironment. High expression of MASP1 was significantly correlated with survival benefits from immune checkpoint inhibitor therapy, while SLC9A3R2 was closely related to the efficacy of PD-L1 inhibitors and CTLA4 inhibitors. CONCLUSIONS: MASP1 and SLC9A3R2, as two key genes that may be related to the immune microenvironment of rectal cancer radiotherapy, deserve further exploration of their roles in the mechanism. The combination of radiotherapy and immunotherapy holds promising prospects in the treatment of rectal cancer, and exploration of related mechanisms will provide new strategies and targets for the treatment of various tumors and rectal cancer.

Bioinformatics↗

MAPping the eukaryotic tree of life: structure, function, and evolution of the MAP215/Dis1 family of microtubule-associated proteins.

The MAP215/Dis1 family of proteins is an evolutionarily ancient family of microtubule-associated proteins, with characterized members in all major kingdoms of eukaryotes, including fungi (Stu2 in S. cerevisiae, Dis1 and Alp14 in S. pombe), Dictyostelium (DdCP224), plants (Mor1 in A. thaliana and TMBP200 in N. tabaccum), and animals (Zyg9 in C. elegans, Msps in Drosophila, XMAP215 in Xenopus, and ch-TOG in humans). All MAP215/Dis1 proteins (with the exception of those in plants) localize to microtubule-organizing centers (MTOCs), including spindle pole bodies in yeast and centrosomes in animals, and all bind to microtubules in vitro and?or in vivo. Diverse roles in regulating microtubule assembly and organization have been proposed for individual family members, and a substantial body of evidence suggests that MAP215/Dis1-related proteins play critical roles in the assembly and function of the meiotic/mitotic spindles and/or cell division. An extensive search of public databases (including both EST and genome databases) identified partial sequences predicted to encode more than three dozen new members of the MAP215/Dis1 family, including putative MAP215/Dis1-related proteins in Giardia lamblia and four other protists, sixteen additional species of fungi, six plants, and twelve animals. The structure and function of MAP215/Dis1 proteins are discussed in relation to the evolution of this ancient family of microtubule-associated proteins.

Amino Acid Sequence↗

The YEASTRACT database: a tool for the analysis of transcription regulatory associations in Saccharomyces cerevisiae.

We present the YEAst Search for Transcriptional Regulators And Consensus Tracking (YEASTRACT; www.yeastract.com) database, a tool for the analysis of transcription regulatory associations in Saccharomyces cerevisiae. This database is a repository of 12 346 regulatory associations between transcription factors and target genes, based on experimental evidence which was spread throughout 861 bibliographic references. It also includes 257 specific DNA-binding sites for more than a hundred characterized transcription factors. Further information about each yeast gene included in the database was obtained from Saccharomyces Genome Database (SGD), Regulatory Sequences Analysis Tools and Gene Ontology (GO) Consortium. Computational tools are also provided to facilitate the exploitation of the gathered data when solving a number of biological questions as exemplified in the Tutorial also available on the system. YEASTRACT allows the identification of documented or potential transcription regulators of a given gene and of documented or potential regulons for each transcription factor. It also renders possible the comparison between DNA motifs, such as those found to be over-represented in the promoter regions of co-regulated genes, and the transcription factor-binding sites described in the literature. The system also provides an useful mechanism for grouping a list of genes (for instance a set of genes with similar expression profiles as revealed by microarray analysis) based on their regulatory associations with known transcription factors.

Binding Sites↗

Contribution of the cyclic nucleotide phosphodiesterases PdeA and PdeB to adaptation of Myxococcus xanthus cells to osmotic or high-temperature stress.

A tBLASTn search of the Myxococcus xanthus genome database at The Institute for Genomic Research (TIGR) identified three genes (pdeA, pdeB, and pdeC) that encode proteins homologous to 3',5'-cyclic nucleotide phosphodiesterase. pdeA, pdeB, and pdeC mutants, constructed by replacing a part of the gene with the kanamycin or tetracycline resistance gene, showed normal growth, development, and germination under nonstress conditions. However, the spores of mutants, especially the pdeA and pdeB mutants, placed under osmotic stress germinated earlier than the wild-type spores. The phenotype was the opposite of that of the receptor-type adenylyl cyclase (cyaA or cyaB) mutant. Also, pdeA and pdeB mutants were found to have impaired growth under the condition of high-temperature stress. Intracellular cyclic AMP (cAMP) levels of pdeA or pdeB mutant cells under these stressful conditions were about 1.3-fold to 2.0-fold higher than those of wild-type cells. These results suggest that PdeA and PdeB may be involved in osmotic adaptation during spore germination and temperature adaptation during vegetative growth through the regulation of cAMP levels.

3',5'-Cyclic-AMP Phosphodiesterases↗

Glycome project: concept, strategy and preliminary application to Caenorhabditis elegans.

Glycans play a central role as potential mediators between complex cell societies, because all living organisms consist of cells covered with diverse carbohydrate chains reflecting various cell types and states. However, we have no idea how diverse these carbohydrate chains actually are. The main purpose of this article is to persuade life scientists to realize the fundamental importance of taking some action by becoming involved in "glycomics". "Glycome" is a term meaning the whole set of glycans produced by individual organisms, as the third bioinformative macromolecules to be elucidated next to the genome and proteome. Here a basic strategy is presented. The essence of the project includes the following: (a) glycopeptides, but not glycans released from their core proteins, are targeted for linkage to genome databases; (b) Caenorhabditis elegans is used as the first model organism for this project, since its genome project has already been completed; (c) four essential attributes are adopted to characterize each glycopeptide: (i) cosmid identification number (ID), (ii) molecular weight (M(r)), (iii) retention (Rs) of pyridylaminated (PA) oligosaccharides in 2-D mapping, and (iv) dissociation constants (Kd's) of PA-oligosaccharides for a set of lectins. Thus, the obtained ID, M(r), R and Kd's construct the glycome database, which will be open as the previous genome and proteome databases. For the project to proceed the "glyco-catch" method is proposed, where a group of target glycopeptides are captured by means of lectin-affinity chromatography after protease digestion. Already glycopeptides from asialofetuin and ovalbumin were successfully captured by galectin-agarose and Con A-agarose, respectively. Further, to examine the practical validity of the method, we extracted membrane proteins from C. elegans with 1% Triton X-100, and isolated specific glycopeptides by use of the same galectin column. One of the glycopeptides was successfully identified in the C. elegans genome database. Finally, for determination of Kd between glycopeptides and lectins, a recently reinforced frontal affinity chromatography (FAC) is proposed as an alternative to define glycan structures in place of determining every covalent structure.

Animals↗

The SUPERFAMILY database in structural genomics.

The SUPERFAMILY hidden Markov model library representing all proteins of known structure predicts the domain architecture of protein sequences and classifies them at the SCOP superfamily level. This analysis has been carried out on all completely sequenced genomes. The ways in which the database can be useful to crystallographers is discussed, in particular with a view to high-throughput structure determination. The application of the SUPERFAMILY database to different target-selection strategies is suggested: novel folds, novel domain combinations and targeted attacks on genomes. Use of the database for more general inquiry in the context of structural studies is also explained. The database provides evolutionary relationships between target proteins and other proteins of known structure through the SCOP database, genome assignments and multiple sequence alignments.

Amino Acid Sequence↗

Bridging the gap between molecular genetics and metabolic medicine: access to genetic information.

UNLABELLED: Thanks to the World Wide Web, most results of research in genetics are made available in public databases. At the present time there are resources on genetic diseases, genes and their location, mutations of already cloned genes and on laboratories performing the mutation analysis. The main resources on phenotypes are On-line Mendelian Inheritance in Man (OMIM), Pedbase, GeneClinics, London Dysmorphology Database (LDDB) and Orphanet. The main resources on human genes are, in addition to OMIM, the Genome Database, Genatlas and Genecard. There are also two major sequence databases. All of them can be queried using the OMIM number of the disease. Central databases of mutations, as well as locus specific databases have been created. Their list is maintained at the Human Genome Organisation mutation database initiative website. Several initiative have been taken to integrate all these data and help the clinician to find out quickly what he/she needs. The website of the National Center for Biotechnology Information is the best example of such an effort with sections on diseases, a genome guide, and locus links. Several databases of genetic testing resources have been established. GeneTests is an on-line genetics resource that contains a directory of North American laboratories providing testing for heritable disorders. Orphanet is a similar database on French services which is in the process of becoming a European database. CONCLUSION: Even if clinicians do not have as many services at their disposal as the molecular geneticists, various useful databases already exist and should no longer be ignored in practice.

Databases, Factual↗

Open reading frame yjbI of Bacillus subtilis codes for truncated hemoglobin.

A hypothetical open reading frame from Bacillus subtilis genome, yjbI [NCBI genome database Accession No. ] having homology to many globin and globin-like proteins from different microbial genomes, was selectively amplified from the chromosomal DNA of B. subtilis strain DB104 based on genome sequence database of B. subtilis strain 168. The gene was cloned and over-expressed in Escherichia coli under the transcriptional control of tandem lambda P(L) and P(R) promoters, and the protein was purified to homogeneity. The single-chain monomeric hemoglobin-like protein is stable to the extent of 5.45 kcal/mol at 25 degrees C, binds carbon mono-oxide, and shows optical spectra characteristic of hemoproteins. The protein also exhibits peroxidase-like activity. This is the first report of a truncated bacterial globin endowed with peroxidase-like activity. The activity is enhanced in the presence of urea and guanidine hydrochloride, more so in the presence of the latter. Presumably, only a small portion of the protein is involved in peroxidase activity, which is exposed with increasing concentration of the denaturants.

Amino Acid Sequence↗

GXD: a gene expression database for the laboratory mouse. The Gene Expression Database Group.

The Gene Expression Database (GXD) is a community resource that stores and integrates expression information for the laboratory mouse, with a particular emphasis on mouse development, and makes these data freely available in formats appropriate for comprehensive analysis. GXD is implemented as a relational database and integrated with the Mouse Genome Database (MGD) to enable global analysis of genotype, expression and phenotype information. Interconnections with sequence databases and with databases from other species further extend GXD's utility for the analysis of gene expression data. GXD is available through the Mouse Genome Informatics Web Site at http://www.informatics.jax.org/

Animals↗

TcruziDB: an integrated, post-genomics community resource for Trypanosoma cruzi.

TcruziDB (http://TcruziDB.org) is an integrated post-genomics database for the parasitic organism, Trypanosoma cruzi, the causative agent of Chagas' disease. TcruziDB was established in 2003 as a flat-file database with tools for mining the unannotated sequence reads and preliminary contig assemblies emerging from the Tri-Tryp genome consortium (TIGR/SBRI/Karolinska). Today, TcruziDB houses the recently published assembled genomic contigs and annotation provided by the genome consortium in a relational database supported by the Genomics Unified Schema (GUS) architecture. The combination of an annotated genome and a relational architecture has facilitated the integration of genomic data with expression data (proteomic and EST) and permitted the construction of automated analysis pipelines. TcruziDB has accepted, and will continue to accept the deposition of genomic and functional genomic datasets contributed by the research community.

Animals↗

GLIDA: GPCR-ligand database for chemical genomic drug discovery.

G-protein coupled receptors (GPCRs) represent one of the most important families of drug targets in pharmaceutical development. GPCR-LIgand DAtabase (GLIDA) is a novel public GPCR-related chemical genomic database that is primarily focused on the correlation of information between GPCRs and their ligands. It provides correlation data between GPCRs and their ligands, along with chemical information on the ligands, as well as access information to the various web databases regarding GPCRs. These data are connected with each other in a relational database, allowing users in the field of GPCR-related drug discovery to easily retrieve such information from either biological or chemical starting points. GLIDA includes structure similarity search functions for the GPCRs and for their ligands. Thus, GLIDA can provide correlation maps linking the searched homologous GPCRs (or ligands) with their ligands (or GPCRs). By analyzing the correlation patterns between GPCRs and ligands, we can gain more detailed knowledge about their interactions and improve drug design efforts by focusing on inferred candidates for GPCR-specific drugs. GLIDA is publicly available at http://gdds.pharm.kyoto-u.ac.jp:8081/glida. We hope that it will prove very useful for chemical genomic research and GPCR-related drug discovery.

Animals↗

Private detection of relatives in forensic genomics using homomorphic encryption.

BACKGROUND: Forensic analysis heavily relies on DNA analysis techniques, notably autosomal Single Nucleotide Polymorphisms (SNPs), to expedite the identification of unknown suspects through genomic database searches. However, the uniqueness of an individual's genome sequence designates it as Personal Identifiable Information (PII), subjecting it to stringent privacy regulations that can impede data access and analysis, as well as restrict the parties allowed to handle the data. Homomorphic Encryption (HE) emerges as a promising solution, enabling the execution of complex functions on encrypted data without the need for decryption. HE not only permits the processing of PII as soon as it is collected and encrypted, such as at a crime scene, but also expands the potential for data processing by multiple entities and artificial intelligence services. METHODS: This study introduces HE-based privacy-preserving methods for SNP DNA analysis, offering a means to compute kinship scores for a set of genome queries while meticulously preserving data privacy. We present three distinct approaches, including one unsupervised and two supervised methods, all of which demonstrated exceptional performance in the iDASH 2023 Track 1 competition. RESULTS: Our HE-based methods can rapidly predict 400 kinship scores from an encrypted database containing 2000 entries within seconds, capitalizing on advanced technologies like Intel AVX vector extensions, Intel HEXL, and Microsoft SEAL HE libraries. Crucially, all three methods achieve remarkable accuracy levels (ranging from 96% to 100%), as evaluated by the auROC score metric, while maintaining robust 128-bit security. These findings underscore the transformative potential of HE in both safeguarding genomic data privacy and streamlining precise DNA analysis. CONCLUSIONS: Results demonstrate that HE-based solutions can be computationally practical to protect genomic privacy during screening of candidate matches for further genealogy analysis in Forensic Genetic Genealogy (FGG).

Humans↗

Beetling around the genome.

The red flour beetle, Tribolium castaneum, has been selected for whole genome shotgun sequencing in the next year. In this minireview, we discuss some of the genetic and genomic tools and biological properties of Tribolium that have established its importance as an organism for agricultural and biomedical research as well as for studies of development and evolution. A Tribolium genomic database, Beetlebase, is being constructed to integrate genetic, genomic and biological data as it becomes available.

Animals↗

Beyond the data deluge: data integration and bio-ontologies.

Biomedical research is increasingly a data-driven science. New technologies support the generation of genome-scale data sets of sequences, sequence variants, transcripts, and proteins; genetic elements underpinning understanding of biomedicine and disease. Information systems designed to manage these data, and the functional insights (biological knowledge) that come from the analysis of these data, are critical to mining large, heterogeneous data sets for new biologically relevant patterns, to generating hypotheses for experimental validation, and ultimately, to building models of how biological systems work. Bio-ontologies have an essential role in supporting two key approaches to effective interpretation of genome-scale data sets: data integration and comparative genomics. To date, bio-ontologies such as the Gene Ontology have been used primarily in community genome databases as structured controlled terminologies and as data aggregators. In this paper we use the Gene Ontology (GO) and the Mouse Genome Informatics (MGI) database as use cases to illustrate the impact of bio-ontologies on data integration and for comparative genomics. Despite the profound impact ontologies are having on the digital categorization of biological knowledge, new biomedical research and the expanding and changing nature of biological information have limited the development of bio-ontologies to support dynamic reasoning for knowledge discovery.

Animals↗

Micado--a network-oriented database for microbial genomes.

MOTIVATION: We created Micado, a database for managing genomic information, as part of the Bacillus subtilis genome programs. Its content will be progressively extended to the whole microbial world. RESULTS: A relational schema is defined for selective queries. It links eubacterial and archaeal sequences, genetic maps for Bacillus subtilis and Escherichia coli, and information on mutants. The latter comes from a new functional analysis project of unknown genes in B subtilis, and the database allows the community to curate information. To help queries from users, a graphical interface is built on SQL access to the database and provided through the WWW. We have automated imports of microbial sequences, and E. coli genetic map, by programming parsers of flat file distributions. These ensure smooth updates from molecular biology repositories on the Internet. Hyperlinks are created as a complement, to reference other general and specialized related information resources.

Bacillus subtilis↗

MBGD: a platform for microbial comparative genomics based on the automated construction of orthologous groups.

The microbial genome database for comparative analysis (MBGD) is a comprehensive platform for microbial comparative genomics. The central function of MBGD is to create orthologous groups among multiple genomes from precomputed all-against-all similarity relationships using the DomClust algorithm. The database now contains >300 published genomes and the number continues to grow. For researchers who are interested in ongoing genome projects, we have now started a new service called 'My MBGD,' which allows users to add their own genome sequences to MBGD for the purpose of identifying orthologs among both the new and the existing genomes. Furthermore, in order to make available the rapidly accumulating information on closely related genome sequences, we enhanced the interface for pairwise genome comparisons using the CGAT interface, which allows users to see nucleotide sequence alignments of non-coding as well as coding regions. MBGD is available at http://mbgd.genome.ad.jp/.

Algorithms↗

PlasmoDB: the Plasmodium genome resource. A database integrating experimental and computational data.

PlasmoDB (http://PlasmoDB.org) is the official database of the Plasmodium falciparum genome sequencing consortium. This resource incorporates the recently completed P. falciparum genome sequence and annotation, as well as draft sequence and annotation emerging from other Plasmodium sequencing projects. PlasmoDB currently houses information from five parasite species and provides tools for intra- and inter-species comparisons. Sequence information is integrated with other genomic-scale data emerging from the Plasmodium research community, including gene expression analysis from EST, SAGE and microarray projects and proteomics studies. The relational schema used to build PlasmoDB, GUS (Genomics Unified Schema) employs a highly structured format to accommodate the diverse data types generated by sequence and expression projects. A variety of tools allow researchers to formulate complex, biologically-based, queries of the database. A stand-alone version of the database is also available on CD-ROM (P. falciparum GenePlot), facilitating access to the data in situations where internet access is difficult (e.g. by malaria researchers working in the field). The goal of PlasmoDB is to facilitate utilization of the vast quantities of genomic-scale data produced by the global malaria research community. The software used to develop PlasmoDB has been used to create a second Apicomplexan parasite genome database, ToxoDB (http://ToxoDB.org).

Animals↗

Genomic evidence for the absence of a functional cholesteryl ester transfer protein gene in mice and rats.

Mice and rats are naturally deficient in cholesteryl ester transfer protein (CETP) activity, although the reason behind the deficiency in activity is unknown. A search of mouse genome databases revealed sequences resembling 7 of the 16 human exons. However, these sequences could not code for a functional CETP. Analysis of the rat genome using Southern blotting revealed sequences complementary to human CETP cDNA, but RNase protection assays were unable to detect any Cetp gene expression in liver, adipose, or muscle. A search of rat whole-genome shotgun databases revealed exon-like sequences that would be unable to code for a functional CETP. An Ap3s1 pseudogene lay immediately upstream of the CETP-like sequences in mouse, but was nearly identical to the functional gene and unlikely to have been inserted prior to mouse-rat divergence. In contrast, a deletion leading to a nonsense codon was found in the exon 11-like sequences of both rat and mouse and not in any other species. Thus, the lack of CETP activity in both the mouse and the rat is most likely due to an evolutionary event that occurred before these species diverged and not to altered regulation of the gene or function of the gene product.

Adaptor Protein Complex 3↗