Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

In silico analysis of 2085 clones from a normalized rat vestibular periphery 3' cDNA library.

The inserts from 2400 cDNA clones isolated from a normalized Rattus norvegicus vestibular periphery cDNA library were sequenced and characterized. The Wackym-Soares vestibular 3' cDNA library was constructed from the saccular and utricular maculae, the ampullae of all three semicircular canals and Scarpa's ganglia containing the somata of the primary afferent neurons, microdissected from 104 male and female rats. The inserts from 2400 randomly selected clones were sequenced from the 5' end. Each sequence was analyzed using the BLAST algorithm compared to the Genbank nonredundant, rat genome, mouse genome and human genome databases to search for high homology alignments. Of the initial 2400 clones, 315 (13%) were found to be of poor quality and did not yield useful information, and therefore were eliminated from the analysis. Of the remaining 2085 sequences, 918 (44%) were found to represent 758 unique genes having useful annotations that were identified in databases within the public domain or in the published literature; these sequences were designated as known characterized sequences. 1141 sequences (55%) aligned with 1011 unique sequences had no useful annotations and were designated as known but uncharacterized sequences. Of the remaining 26 sequences (1%), 24 aligned with rat genomic sequences, but none matched previously described rat expressed sequence tags or mRNAs. No significant alignment to the rat or human genomic sequences could be found for the remaining 2 sequences. Of the 2085 sequences analyzed, 86% were singletons. The known, characterized sequences were analyzed with the FatiGO online data-mining tool (http://fatigo.bioinfo.cnio.es/) to identify level 5 biological process gene ontology (GO) terms for each alignment and to group alignments with similar or identical GO terms. Numerous genes were identified that have not been previously shown to be expressed in the vestibular system. Further characterization of the novel cDNA sequences may lead to the identification of genes with vestibular-specific functions. Continued analysis of the rat vestibular periphery transcriptome should provide new insights into vestibular function and generate new hypotheses. Physiological studies are necessary to further elucidate the roles of the identified genes and novel sequences in vestibular function.

Afferent Pathways↗

ZBTB16-associated NK cell alterations reveal shared immunometabolic signatures linking primary Sjögren's syndrome and type 1 diabetes mellitus.

BACKGROUND: Primary Sjögren's syndrome (pSS) and type 1 diabetes mellitus (T1DM) share immune-inflammatory features, yet conserved pathogenic signatures linking these autoimmune disorders remain incompletely understood. The present research sought to uncover common molecular markers and dissect the underlying immune-metabolic cross-talk underlying pSS and T1DM. METHODS: Gene expression profiles of patients with pSS and T1DM were retrieved from the Gene Expression Omnibus database, normalized, and corrected for batch effects prior to downstream analyses. Overlapping potential biomarkers were screened by integrating differential expression analysis, weighted gene co-expression network analysis and least absolute shrinkage and selection operator regression. Functional enrichment based on Gene Ontology and Kyoto Encyclopedia of Genes and Genomes databases was implemented to interpret gene biological properties, and a protein-protein interaction network was further established afterwards. Diagnostic performance was evaluated using receiver operating characteristic analysis. Experimental validation was conducted in non-obese diabetic (NOD) mice using quantitative PCR, immunohistochemistry, and flow cytometry. The CIBERSORT algorithm was adopted to quantify immune cell infiltration levels. RESULTS: ZBTB16 was identified as a shared hub biomarker in both pSS and T1DM and exhibited favorable diagnostic performance. Experimental validation confirmed significantly reduced ZBTB16 expression in peripheral blood mononuclear cells, salivary gland tissues, and pancreatic tissues of NOD mice. Gene Set Enrichment Analysis indicated that ZBTB16-associated signatures were enriched in mitochondrial-related processes, neuroactive ligand-receptor interactions, and ribosome-related pathways. Immune infiltration analysis revealed that resting natural killer (NK) cells were positively correlated with ZBTB16 expression in both diseases. Flow cytometric analysis further confirmed a reduced proportion of resting NK cells in peripheral blood of NOD mice, consistent with the CIBERSORT-based prediction. CONCLUSION: This study identifies ZBTB16 as a shared biomarker linking pSS and T1DM. Reduced resting NK-cell abundance was consistently observed in both computational and experimental analyses, and bioinformatic correlation analysis suggested a positive association with ZBTB16 expression. These findings provide evidence for shared molecular and immunological signatures underlying the two autoimmune disorders and support further investigation of the biological role and diagnostic value of ZBTB16 in pSS and T1DM.

Sjogren's Syndrome↗

Evidence supporting predicted metabolic pathways for Vibrio cholerae: gene expression data and clinical tests.

Vibrio cholerae, the etiological agent of the diarrheal illness cholera, can kill an infected adult in 24 h. V.cholerae lives as an autochthonous microbe in estuaries, rivers and coastal waters. A better understanding of its metabolic pathways will assist the development of more effective treatments and will provide a deeper understanding of how this bacterium persists in natural aquatic habitats. Using the completed V.cholerae genome sequence and PathoLogic software, we created VchoCyc, a pathway-genome database that predicted 171 likely metabolic pathways in the bacterium. We report here experimental evidence supporting the computationally predicted pathways. The evidence comes from microarray gene expression studies of V.cholerae in the stools of three cholera patients [D. S. Merrell, S. M. Butler, F. Qadri, N. A. Dolganov, A. Alam, M. B. Cohen, S. B. Calderwood, G. K. Schoolnik and A. Camilli (2002) Nature, 417, 642-645.], from gene expression studies in minimal growth conditions and LB rich medium, and from clinical tests that identify V.cholerae. Expression data provide evidence supporting 92 (53%) of the 171 pathways. The clinical tests provide evidence supporting seven pathways, with six pathways supported by both methods. VchoCyc provides biologists with a useful tool for analyzing this organism's metabolic and genomic information, which could lead to potential insights into new anti-bacterial agents. VchoCyc is available in the BioCyc database collection (http://BioCyc.org).

Bacteriological Techniques↗

Bioinformatic and experimental tools for identification of single-nucleotide polymorphisms in genes with a potential role for the development of the insulin resistance syndrome.

OBJECTIVES: Genes with a possible role for the development of the insulin resistance syndrome (IRS) were scanned for novel single-nucleotide polymorphisms (SNPs) using bioinformatics. METHODS: GenBank mRNA sequences were compared to the human EST database using gapped BLAST, software that is available on the internet. Mismatches between the search and the EST sequences indicated potential SNPs. Thirty-two SNPs in 13 genes were randomly chosen for experimental verification. PCR and direct sequencing were used to determine the 'true' SNPs. A random sample of 30 Swedish men with slightly elevated diastolic blood pressure (85-94 mmHg) obtained from a population-based study was selected for the sequencing. After completion of these stages, the potential SNPs were checked against the large and rapidly expanding SNP databases HGBASE and NCBI. RESULTS: EST searches of 146 genes revealed 106 potential SNPs in 44 genes. Experimental analysis of 32 of these potential SNPs verified two SNPs; endothelin receptor A 1471 G/C (3' UTR) and PAI-1 Trp514Arg from a T/C exchange. These two SNPs were also identified in the NCBI and HGBASE databases together with two polymorphisms that were not experimentally identified in our homogeneous Swedish population. Overall, the HGBASE and NCBI databases contained entries of 22% (23 out of 106) of the SNPs identified through our EST searches. CONCLUSIONS: In the search for genetic variations causing complex diseases like IRS in homogeneous populations (such as the Swedish one used here), important information can be obtained through bioinformatic searches of human genome databases and experimental verification.

Adult↗

Semiparametric efficient estimation of small genetic effects in large-scale population cohorts.

Population genetics seeks to quantify DNA variant associations with traits or diseases, as well as interactions among variants and with environmental factors. Computing millions of estimates in large cohorts in which small effect sizes and tight confidence intervals are expected, necessitates minimizing model-misspecification bias to increase power and control false discoveries. We present TarGene, a unified statistical workflow for the semi-parametric efficient and double robust estimation of genetic effects including $ k $-point interactions among categorical variables in the presence of confounding and weak population dependence. $ k $-point interactions, or Average Interaction Effects (AIEs), are a direct generalization of the usual average treatment effect (ATE). We estimate genetic effects with cross-validated and/or weighted versions of Targeted Minimum Loss-based Estimators (TMLE) and One-Step Estimators (OSE). The effect of dependence among data units on variance estimates is corrected by using sieve plateau variance estimators based on genetic relatedness across the units. We present extensive realistic simulations to demonstrate power, coverage, and control of type I error. Our motivating application is the targeted estimation of genetic effects on trait, including two-point and higher-order gene-gene and gene-environment interactions, in large-scale genomic databases such as UK Biobank and All of Us. All cross-validated and/or weighted TMLE and OSE for the AIE $ k $-point interaction, as well as ATEs, conditional ATEs and functions thereof, are implemented in the general purpose Julia package TMLE.jl. For high-throughput applications in population genomics, we provide the open-source Nextflow pipeline and software TarGene which integrates seamlessly with modern high-performance and cloud computing platforms.

Humans↗

D4 Dopamine receptor genes of zebrafish and effects of the antipsychotic clozapine on larval swimming behaviour.

Zebrafish, a model developmental genetic organism, is being increasingly used in behavioural studies. We have initiated studies designed to evaluate the response of zebrafish to antipsychotic drugs. This study focuses on characterization of zebrafish D4 dopamine receptors (D4Rs) and the response of larval zebrafish to the atypical antipsychotic clozapine. The D4R is of interest because of its high affinity for clozapine, while interest in clozapine stems from its effectiveness in reducing symptoms in acutely psychotic, treatment-resistant schizophrenic patients. By mining the zebrafish genomic database, we identified three distinct D4R genes, drd4a, drd4b and drd4c, and generated full-length open reading frames encoding each of the three D4Rs by reverse transcription-polymerase chain reaction. Gene mapping studies showed that each D4R gene mapped to a distinct chromosomal location in the zebrafish genome, and each gene exhibited a unique expression profile during embryogenesis. When administered to larval zebrafish, clozapine produced a rapid and profound effect on locomotor activity. The effect of clozapine was dose-dependent, resulted in hypoactivity and was prevented by the D4-selective agonist ABT-724. Our data suggest that the inhibitory effect of clozapine on the locomotor activity of larval zebrafish may be mediated through D4Rs.

Amino Acid Sequence↗

Analysis of reading frame and expressional regulation of randomly selected promoter-proximal genes in Escherichia coli.

The expression of seventy-seven randomly cloned genes of Escherichia coli was examined following a variety of treatments including heat shock, glucose starvation, phosphate starvation, ammonium starvation or osmotic shock, with the aid of lacZ reporter gene protein fusions on multicopy plasmids. Two of 77 genes (amr and yigL) had not previously been identified as protein encoding open-reading frames (ORFs) in annotations of the E. coli genome database. Thirteen genes exhibited significant changes in expression in response to at least one of the treatments, and six of them appeared to be controlled by more than one sigma (sigma) factor of RNA polymerase. This study thus allows us not only to identify the reading frame of the genomic genes but also to support the hypothesis earlier proposed that a significant proportion of genes in E. coli are involved in adaptations to various stresses to which the organism is likely to be exposed in the environment.

Amino Acid Sequence↗

An efficient delivery of historical information for the Mendelian Inheritance in Man database.

The ability to manage information with regard to changes in a database is critical for quality control. This information can also provide audit trails about the time of the change and the person who made the change. In addition, historical information can provide the proper context in which to interpret the relationships between the current and past data. In most genomic databases, only the most recent copy of the information is presented to the user, thereby losing the audit trail and the historical context. Therefore, we have constructed a delivery mechanism for the historical information in the Mendelian Inheritance in Man database. Furthermore, this feature was designed to optionally display only the changes so that the user can bypass the unchanged portions of the text. It was anticipated that technical problems would influence the acceptance of this information delivery. However, the involvement of the editorial staff became the critical factor.

Database Management Systems↗

Genomic BLAST: custom-defined virtual databases for complete and unfinished genomes.

BLAST (Basic Local Alignment Search Tool) searches against DNA and protein sequence databases have become an indispensable tool for biomedical research. The proliferation of the genome sequencing projects is steadily increasing the fraction of genome-derived sequences in the public databases and their importance as a public resource. We report here the availability of Genomic BLAST, a novel graphical tool for simplifying BLAST searches against complete and unfinished genome sequences. This tool allows the user to compare the query sequence against a virtual database of DNA and/or protein sequences from a selected group of organisms with finished or unfinished genomes. The organisms for such a database can be selected using either a graphic taxonomy-based tree or an alphabetical list of organism-specific sequences. The first option is designed to help explore the evolutionary relationships among organisms within a certain taxonomy group when performing BLAST searches. The use of an alphabetical list allows the user to perform a more elaborate set of selections, assembling any given number of organism-specific databases from unfinished or complete genomes. This tool, available at the NCBI web site http://www.ncbi.nlm.nih.gov/cgi-bin/Entrez/genom_table_cgi, currently provides access to over 170 bacterial and archaeal genomes and over 40 eukaryotic genomes.

Amino Acid Sequence↗

Evolutionary pattern of the gtwin retrotransposon in the Drosophila melanogaster subgroup.

The gtwin retrotransposon was recently discovered in the Drosophila melanogaster genome and it is evolutionarily closer to gypsy endogenous retrovirus. This study has identified gtwin homologous sequences in the genome of D. simulans, D. sechellia, D. erecta and D. yakuba by performing homology searches against the public genome database of Drosophila species. The phylogenetic analyses of the gtwin env gene sequences of these species have shown some incongruities with the host species phylogeny, suggesting some horizontal transfer events for this retroelement. Moreover, we reported the existence of DNA sequences putatively encoding full-length Env proteins in the genomes of Drosophila species other than D. melanogaster. The results suggest that the gtwin element may be an infectious retrovirus able to invade the genome of new species, supporting the gtwin evolutionary picture shown in this work.

Animals↗

XET activity is found near sites of growth and cell elongation in bryophytes and some green algae: new insights into the evolution of primary cell wall elongation.

BACKGROUND AND AIMS: In angiosperms xyloglucan endotransglucosylase (XET)/hydrolase (XTH) is involved in reorganization of the cell wall during growth and development. The location of oligo-xyloglucan transglucosylation activity and the presence of XTH expressed sequence tags (ESTs) in the earliest diverging extant plants, i.e. in bryophytes and algae, down to the Phaeophyta was examined. The results provide information on the presence of an XET growth mechanism in bryophytes and algae and contribute to the understanding of the evolution of cell wall elongation in general. METHODS: Representatives of the different plant lineages were pressed onto an XET test paper and assayed. XET or XET-related activity was visualized as the incorporation of fluorescent signal. The Physcomitrella genome database was screened for the presence of XTHs. In addition, using the 3' RACE technique searches were made for the presence of possible XTH ESTs in the Charophyta. KEY RESULTS: XET activity was found in the three major divisions of bryophytes at sites corresponding to growing regions. In the Physcomitrella genome two putative XTH-encoding cDNA sequences were identified that contain all domains crucial for XET activity. Furthermore, XET activity was located at the sites of growth in Chara (Charophyta) and Ulva (Chlorophyta) and a putative XTH ancestral enzyme in Chara was identified. No XET activity was identified in the Rhodophyta or Phaeophyta. CONCLUSIONS: XET activity was shown to be present in all major groups of green plants. These data suggest that an XET-related growth mechanism originated before the evolutionary divergence of the Chlorobionta and open new insights in the evolution of the mechanisms of primary cell wall expansion.

Amino Acid Sequence↗

Analysis of Arabidopsis genome sequence reveals a large new gene family in plants.

A detailed analysis of the currently available Arabidopsis thaliana genomic sequence has revealed the presence of a large number of open reading frames with homology to the stigmatic self-incompatibility (S) genes of Papaver rhoeas. The products of these potential genes are all predicted to be relatively small, basic, secreted proteins with similar predicted secondary structures. We have named these potential genes SPH (S-protein homologues). Their presence appears to have been largely missed by the prediction methods currently used on the genomic sequence. Equivalent homologues could not be detected in the human, microbial, Drosophila or C. elegans genomic databases, suggesting a function specific to plants. Preliminary RT-PCR analysis indicates that at least two members of the family (SPH1, SPH8) are expressed, with expression being greatest in floral tissues. The gene family may total more than 100 members, and its discovery not only illustrates the importance of the genome sequencing efforts, but also indicates the extent of information which remains hidden after the initial trawl for potential genes.

Arabidopsis↗

SPEN inactivation drives resistance to androgen receptor pathway inhibitors in metastatic prostate cancer.

PURPOSE: Treatment intensification with androgen receptor pathway inhibitors (ARPIs) has become the standard of care for patients with metastatic prostate cancer. However, there remains an unmet need to identify biomarkers for treatment resistance. Here, we identify SPEN inactivation as a driver of ARPI resistance. EXPERIMENTAL DESIGN: Pre-clinical studies were performed in LNCaP and VCaP cell lines. Data from a nationwide prostate cancer clinico-genomic database were extracted. Log-rank test and Cox proportional hazards models were used to compare time to next treatment (TTNT) on ARPI with/without SPEN mutations. SPEN immunohistochemistry was performed on a rapid autopsy metastatic tissue microarray. RESULTS: SPEN was identified as a top enzalutamide resistance hit in an unbiased genome-wide loss-of-function screen. SPEN inactivation results in upregulation of cell cycle proliferation and basal/stem cell activity as well as increased translation of pro-oncogenic genes. In a large patient cohort (N=6828), SPEN mutations are enriched following treatment with ARPIs (2.1% to 3.6%, p=0.001) and correlate with shorter TTNT on ARPI in patients with metastatic hormone-sensitive prostate cancer (6.4 vs 29.7 months, HR 2.67, p=0.02). In a metastatic rapid autopsy cohort (N=181), low SPEN H-score is associated with shorter time on abiraterone (5.0 vs 7.9 months, p=0.023) in metastatic castration-resistant prostate cancer. CONCLUSIONS: In real-world cohorts, loss of SPEN function across genomic, transcriptomic, and protein levels is associated with reduced benefit from ARPI therapy in metastatic prostate cancer. These findings identify SPEN inactivation as a clinically relevant biomarker of ARPI resistance that warrants prospective evaluation to guide treatment selection.

Journal Article↗

There exist at least 30 human G-protein-coupled receptors with long Ser/Thr-rich N-termini.

We report six novel members of the superfamily of human G-protein coupled receptors (GPCRs) found by searches in the human genome databases, termed GPR123, GPR124, GPR125, GPR126, GPR127, and GPR128. Phylogenetic analysis demonstrates that these are additional members of the family of GPCRs with long N-termini, previously termed EGF-7TM, LNB-7TM, B2 or LN-7TM, showing that there exist at least 30 such GPCRs in the human genome. Three of these receptors form their own phylogenetic cluster, while two other places in a cluster with the previously reported HE6 and GPR56 (TM7XN1) and one with EMR1-3. All the novel receptors have a GPS domain in their N-terminus, except GPR123, as well as long Ser/Thr rich regions forming mucin-like stalks. GPR124 and GPR125 have a leucine rich repeat (LRR), an immunoglobulin (Ig) domain, and a hormone-binding domain (HBD). The Ig domain shows similarities to motilin and titin, while the LRR domain shows similarities to LRIG1 and SLIT1-2. GPR127 has one EGF domain while GPR126 and GPR128 do not contain domains that are readily recognized in other proteins beyond the GPS domain. We found several human EST sequences for most of the receptors showing differential expression patterns, which may indicate that some of these receptors participate in central functions while others are more likely to have a role in the immune or reproductive systems.

Amino Acid Sequence↗

TestisBank: an internet-based gene sequence database of the testis.

The testis is a highly transcriptionally active organ with hundreds of genes expressed during different stages of spermatogenesis. Scientists working on the testis are restricted to using two sources for further information on testicular genes, GenBank and MEDLINE. However, these two databases are not completely linked and give only very little information on the cellular type of expression or whether a gene is cloned in other tissues but is also expressed in the testis. We have generated an organ-specific database named TestisBank, which is capable of retrieving a more complete set of genes expressed in the different cell types of the testis. Furthermore, it extends to the epididymis which plays an important role in germ cell maturation. TestisBank is automatically updated to match the current status of sequence entries in the genome databases and it provides the user with an interface convenient to handle and provides useful links related to male reproduction. The TestisBank is publicly available at http://medweb.uni-muenster.de/TestisBank/. The design of the TestisBank may provide a simple model for the development of databases in which molecular and literature data are merged, thereby allowing detailed insight into the contribution of different cells/genes to the complex process of spermatogenesis.

Databases, Bibliographic↗

Assessment of the relative number of copies of the gene encoding human neutrophil antigen-2a(HNA-2a), CD177, and a homologous pseudogene by quantitative real-time PCR.

Human neutrophil antigen-2a (HNA-2a; NB1) is located on the 58-64 kD NB1 glycoprotein (GP) and is encoded by the gene CD177. Searches of human genome databases have revealed that a pseudogene highly homologous to exons 4-9 of CD177 is located adjacent to CD177 on chromosome 19. The purpose of this study was to document the presence of the pseudogene and determine whether the polymorphic expression of NB1 GP is due to CD177 gene deletions and duplications. Genomic DNA was isolated from leukocytes of 12 subjects. The number of copies of exon 2 of CD177, an exon that is unique to this gene, and the number of copies of exon 9, an exon that is found in both CD177 and the pseudogene, was assessed with quantitative real-time PCR. The ratio of the number of copies of sequences homologous to CD177 exon 9 to the number of copies of exon 2 was 1.5 or greater in 7 of the 12 subjects, suggesting that both CD177 and the homologous pseudogene were present. The ratio of exon 9 to exon 2 in the other 5 subjects ranged from 1 to 1.25, suggesting that the pseudogene was not present in these subjects. However, results of assays were variable and we could not exclude the possibility that all subjects carried the pseudogene. These studies confirmed the presence of the pseudogene homologous to CD177, but quantitative real-time PCR was not precise enough to detect CD177 duplications or deletions.

Journal Article↗

Addressing protein localization within the nucleus.

Bridging the gap between the number of gene sequences in databases and the number of gene products that have been functionally characterized in any way is a major challenge for biology. A key characteristic of proteins, which can begin to elucidate their possible functions, is their subcellular location. A number of experimental approaches can reveal the subcellular localization of proteins in mammalian cells. However, genome databases now contain predicted sequences for a large number of potentially novel proteins that have yet to be studied in any way, let alone have their subcellular localization determined. Here we ask whether using bioinformatics tools to analyse the sequence of proteins whose subnuclear localizations have been determined can reveal characteristics or signatures that might allow us to predict localization for novel protein sequences.

Amino Acid Motifs↗