Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43Linked to original sources

The TIGR Plant Transcript Assemblies database.

The TIGR Plant Transcript Assemblies (TA) database (http://plantta.tigr.org) uses expressed sequences collected from the NCBI GenBank Nucleotide database for the construction of transcript assemblies. The sequences collected include expressed sequence tags (ESTs) and full-length and partial cDNAs, but exclude computationally predicted gene sequences. The TA database includes all plant species for which more than 1000 EST or cDNA sequences are publicly available. The EST and cDNA sequences are first clustered based on an all-versus-all pairwise sequence comparison, followed by the generation of consensus sequences (TAs) from individual clusters. The clustering and assembly procedures use the TGICL tool, Megablast and the CAP3 assembler. The UniProt Reference Clusters (UniRef100) protein database is used as the reference database for the functional annotation of the assemblies. The transcription orientation of each TA is determined based on the orientation of the alignment with the best protein hit. The TA sequences and annotation are available via web interfaces and FTP downloads. Assemblies can be retrieved by a text-based keyword search or a sequence-based BLAST search. The current version of the TA database is Release 2 (July 17, 2006) and includes a total of 215 plant species.

DNA, Complementary↗

TOPOFIT-DB, a database of protein structural alignments based on the TOPOFIT method.

TOPOFIT-DB (T-DB) is a public web-based database of protein structural alignments based on the TOPOFIT method, providing a comprehensive resource for comparative analysis of protein structure families. The TOPOFIT method is based on the discovery of a saturation point on the alignment curve (topomax point) which presents an ability to objectively identify a border between common and variable parts in a protein structural family, providing additional insight into protein comparison and functional annotation. TOPOFIT also effectively detects non-sequential relations between protein structures. T-DB provides users with the convenient ability to retrieve and analyze structural neighbors for a protein; do one-to-all calculation of a user provided structure against the entire current PDB release with T-Server, and pair-wise comparison using the TOPOFIT method through the T-Pair web page. All outputs are reported in various web-based tables and graphics, with automated viewing of the structure-sequence alignments in the Friend software package for complete, detailed analysis. T-DB presents researchers with the opportunity for comprehensive studies of the variability in proteins and is publicly available at http://mozart.bio.neu.edu/topofit/index.php.

Databases, Protein↗

ForestTreeDB: a database dedicated to the mining of tree transcriptomes.

ForestTreeDB is intended as a resource that centralizes large-scale expressed sequence tag (EST) sequencing results from several tree species (http://foresttree.org/ftdb). It currently encompasses 344,878 quality sequences from 68 libraries, from diverse organs of conifer and hybrid poplar trees. It utilizes the Nimbus data model to provide a hosting system for multiple projects, and uses object-relational mapping APIs in Java and Perl for data accesses within an Oracle database designed to be scalable, maintainable and extendable. Transcriptome builds or unigene sets occupy the focal point of the system. Several of the five current species-specific unigenes were used to design microarrays and SNP resources. The ForestTreeDB web application provides the means for multiple combination database queries. It presents the user with a list of discrete queries to retrieve and download large EST datasets or sequences from precompiled unigene assemblies. Functional annotation assignment is not trivial in conifers which are distantly related to angiosperm model plants. Optimal annotations are achieved through database queries that integrate results from several procedures based open-source tools. ForestTreeDB aims to facilitate sequence mining of coherent annotations in multiple species to support comparative genomic approaches. We plan to continuously enrich ForestTreeDB with other resources through collaborations with other genomic projects.

Databases, Nucleic Acid↗

The SUPERFAMILY database in 2007: families and functions.

The SUPERFAMILY database provides protein domain assignments, at the SCOP 'superfamily' level, for the predicted protein sequences in over 400 completed genomes. A superfamily groups together domains of different families which have a common evolutionary ancestor based on structural, functional and sequence data. SUPERFAMILY domain assignments are generated using an expert curated set of profile hidden Markov models. All models and structural assignments are available for browsing and download from http://supfam.org. The web interface includes services such as domain architectures and alignment details for all protein assignments, searchable domain combinations, domain occurrence network visualization, detection of over- or under-represented superfamilies for a given genome by comparison with other genomes, assignment of manually submitted sequences and keyword searches. In this update we describe the SUPERFAMILY database and outline two major developments: (i) incorporation of family level assignments and (ii) a superfamily-level functional annotation. The SUPERFAMILY database can be used for general protein evolution and superfamily-specific studies, genomic annotation, and structural genomics target suggestion and assessment.

Databases, Protein↗

The Universal Protein Resource (UniProt).

The ability to store and interconnect all available information on proteins is crucial to modern biological research. Accordingly, the Universal Protein Resource (UniProt) plays an increasingly important role by providing a stable, comprehensive, freely accessible central resource on protein sequences and functional annotation. UniProt is produced by the UniProt Consortium, formed in 2002 by the European Bioinformatics Institute (EBI), the Protein Information Resource (PIR) and the Swiss Institute of Bioinformatics (SIB). The core activities include manual curation of protein sequences assisted by computational analysis, sequence archiving, development of a user-friendly UniProt web site and the provision of additional value-added information through cross-references to other databases. UniProt is comprised of three major components, each optimized for different uses: the UniProt Archive, the UniProt Knowledgebase and the UniProt Reference Clusters. An additional component consisting of metagenomic and environmental sequences has recently been added to UniProt to ensure availability of such sequences in a timely fashion. UniProt is updated and distributed on a bi-weekly basis and can be accessed online for searches or download at http://www.uniprot.org.

Amino Acid Sequence↗

The CATH domain structure database: new protocols and classification levels give a more comprehensive resource for exploring evolution.

We report the latest release (version 3.0) of the CATH protein domain database (http://www.cathdb.info). There has been a 20% increase in the number of structural domains classified in CATH, up to 86 151 domains. Release 3.0 comprises 1110 fold groups and 2147 homologous superfamilies. To cope with the increases in diverse structural homologues being determined by the structural genomics initiatives, more sensitive methods have been developed for identifying boundaries in multi-domain proteins and for recognising homologues. The CATH classification update is now being driven by an integrated pipeline that links these automated procedures with validation steps, that have been made easier by the provision of information rich web pages summarising comparison scores and relevant links to external sites for each domain being classified. An analysis of the population of domains in the CATH hierarchy and several domain characteristics are presented for version 3.0. We also report an update of the CATH Dictionary of homologous structures (CATH-DHS) which now contains multiple structural alignments, consensus information and functional annotations for 1459 well populated superfamilies in CATH. CATH is directly linked to the Gene3D database which is a projection of CATH structural data onto approximately 2 million sequences in completed genomes and UniProt.

Classification↗

Glycoengineering of cyanobacterial thylakoid membranes for future studies on the role of glycolipids in photosynthesis.

The lipid composition of thylakoid membranes is conserved from cyanobacteria to angiosperms. The predominating components are monogalactosyl- and digalactosyldiacylglycerol. In cyanobacteria, thylakoid membrane biosynthesis starts with the formation of monoglucosyldiacylglycerol which is C4-epimerized to the corresponding galactolipid, whereas in plastids monogalactosyldiacylglycerol is formed at the beginning. This suggests that galactolipids have specific functions in thylakoids. We wanted to investigate whether galactolipids can be replaced by glycosyldiacylglycerols with headgroups differing in their epimeric and anomeric details as well as the attachment point of the terminal hexose in diglycosyldiacylglycerols. For this purpose putative glycosyltransferase sequences were identified in databases to be used for functional expression in various host organisms. From 18 newly identified sequences, four turned out to encode glycosyltransferases catalyzing final steps in glycolipid biosynthesis: two alpha-glucosyltransferases, one beta-galactosyltransferase and one beta-glucosyltransferase. Their functional annotation was based on detailed structural characterization of the new glycolipids formed in the transformant hosts as well as on in vitro enzymatic assays. The expression of alpha-glucosyltransferases in the cyanobacterium Synechococcus resulted in the accumulation of the new alpha-galactosyldiacylglycerol which is ascribed to epimerization of the corresponding glucolipid. The expression of the beta-glucosyltransferase led to a high proportion of new beta-glucosyl-(1-->6)-beta-galactosyldiacylglycerol almost entirely replacing the native digalactosyldiacylglycerol. These results demonstrate that modifications of the glycolipid pattern in thylakoids are possible.

Cloning, Molecular↗

A genome-wide cross-trait analysis characterizes the shared genetic architecture between rheumatoid arthritis and psychiatric disorders.

OBJECTIVES: Patients with RA have a 2- to 3-fold elevated risk of psychiatric disorders, suggesting an underlying genetic link between these phenotypes. However, the shared genetic architectures and pathological mechanisms driving RA-psychiatric disorder comorbidity remain to be fully elucidated. Herein, we performed cross-trait analysis to investigate the shared genetic architecture between RA and psychiatric disorders. METHODS: Leveraging European-ancestry genome-wide association studies (GWASs) datasets of RA (n = 1 026 690) and 10 major psychiatric disorders (n = 14 307-1 222 882), we performed cross-trait pleiotropic analysis to identify the shared pleiotropic loci and genes between RA and psychiatric disorders, followed by functional annotation and Mendelian randomization analysis to explore the pathological mechanisms underlying RA-psychiatric disorder comorbidity. RESULTS: Our analysis revealed significant positive genetic correlations between RA and seven psychiatric disorders, such as major depressive disorder. From these correlations, we identified 61 pleiotropic loci jointly influencing RA and psychiatric disorder risk, along with 208 pleiotropic genes predominantly involved in immune and inflammatory response biological processes. Druggable target exploration identified 21 drug-gene interactions involving pleiotropic genes, with two genes (RHOA and TRAF3) classified in the clinically actionable category, representing potential therapeutic targets for both RA and psychiatric disorders. Mendelian randomization further demonstrated a bidirectional causal relationship between RA and schizophrenia, while supporting the causal roles of attention-deficit/hyperactivity disorder, major depressive disorder and post-traumatic stress disorder in increasing RA risk. CONCLUSION: Our findings elucidate the shared genetic architecture between RA and psychiatric disorders, providing novel insights into the pathological mechanisms underlying their comorbidity and laying the groundwork for improved comorbidity management.

Arthritis, Rheumatoid↗

Comparative toxicogenomic analysis of the hepatotoxic effects of TCDD in Sprague Dawley rats and C57BL/6 mice.

In an effort to further characterize conserved and species-specific mechanisms of 2,3,7,8-tetrachlorodibenzo-p-dioxin (TCDD)-mediated toxicity, comparative temporal and dose-response microarray analyses were performed on hepatic tissue from immature, ovariectomized Sprague Dawley rats and C57BL/6 mice. For temporal studies, rats and mice were gavaged with 10 or 30 microg/kg of TCDD, respectively, and sacrificed after 2, 4, 8, 12, 18, 24, 72, or 168 h while dose-response studies were performed at 24 h. Hepatic gene expression profiles were monitored using custom cDNA microarrays containing 8567 (rat) or 13,361 (mouse) cDNA clones. Affymetrix data from male rats treated with 40 microg/kg TCDD were also included to expand the species comparison. In total, 3087 orthologous genes were represented in the cross-species comparison. Comparative analysis identified 33 orthologous genes that were commonly regulated by TCDD as well as 185 rat-specific and 225 mouse-specific responses. Functional annotation using Gene Ontology identified conserved gene responses associated with xenobiotic/chemical stress and amino acid and lipid metabolism. Rat-specific gene expression responses were associated with cellular growth and lipid metabolism while mouse-specific responses were associated with lipid uptake/metabolism and immune responses. The common and species-specific gene expression responses were also consistent with complementary histopathology, clinical chemistry, hepatic lipid analyses, and reports in the literature. These data expand our understanding of TCDD-mediated gene expression responses and indicate that species-specific toxicity may be mediated by differences in gene expression which may help explain the wide range of species sensitivities and will have important implications in risk assessment strategies.

Animals↗

Expression profiling soybean response to Pseudomonas syringae reveals new defense-related genes and rapid HR-specific downregulation of photosynthesis.

Transcript profiling during susceptible (S) and hypersensitive response-associated resistance (R) interactions was determined in soybean (Glycine max). Pseudomonas syringae pv. glycinea carrying or lacking the avirulence gene avrB, was infiltrated into cultivar Williams 82. Leaf RNA was sampled at 2, 8, and 24 h postinoculation (hpi). Significant changes in transcript abundance were observed for 3,897 genes during the experiment at P < or = 0.000005. Many of the genes showed a similar direction of increase or decrease in abundance in both the S and R responses, but the R response generally showed a significantly greater degree of differential expression. More than 25% of these responsive genes had not been previously reported as being associated with pathogen interactions, as 704 had no functional annotation and 378 had no homolog in National Center for Biotechnology Information databases. The highest number of transcriptional changes was noted at 8 hpi, including the downregulation of 94 chloroplast-associated genes specific to the R response. Photosynthetic measurements were consistent with an R-specific reduction in photosystem II operating efficiency (phiPSII) that was apparent at 8 hpi for the R response with little effect in the S or control treatments. Imaging analyses suggest that the decreased phiPSII was a result of physical damage to PSII reaction centers.

Analysis of Variance↗

Transcriptome analysis of the barley-Fusarium graminearum interaction.

Fusarium head blight (FHB) of barley (Hordeum vulgare L.) is caused by Fusarium graminearum. FHB causes yield losses and reduction in grain quality primarily due to the accumulation of trichothecene mycotoxins such as deoxynivalenol (DON). To develop an understanding of the barley-F. graminearum interaction, we examined the relationship among the infection process, DON concentration, and host transcript accumulation for 22,439 genes in spikes from the susceptible cv. Morex from 0 to 144 h after F. graminearum and water control inoculation. We detected 467 differentially accumulating barley gene transcripts in the F. graminearum-treated plants compared with the water control-treated plants. Functional annotation of the transcripts revealed a variety of infection-induced host genes encoding defense response proteins, oxidative burst-associated enzymes, and phenylpropanoid pathway enzymes. Of particular interest was the induction of transcripts encoding potential trichothecene catabolic enzymes and transporters, and the induction of the tryptophan biosynthetic and catabolic pathway enzymes. Our results define three stages of E graminearum infection. An early stage, between 0 and 48 h after inoculation (hai), exhibited limited fungal development, low DON accumulation, and little change in the transcript accumulation status. An intermediate stage, between 48 and 96 hai, showed increased fungal development and active infection, higher DON accumulation, and increased transcript accumulation. A majority of the host gene transcripts were detected by 72 hai, suggesting that this is an important timepoint for the barley-F. graminearum interaction. A late stage also identified between 96 and 144 hai, exhibiting development of hyphal mats, high DON accumulation, and a reduction in the number of transcripts observed. Our study provides a baseline and hypothesis-generating dataset in barley during F. graminearum infection and in other grasses during pathogen infection.

Fusarium↗

Integrative analysis of cancer-related data using CAP.

The development of human cancer is a highly complex process and can be considered the result of several combined events, such as genetic alterations, disturbance of signal transduction, or failure of immunological surveillance. Cancer-related databases usually focus on specific fields of research, e.g., cancer genetics or cancer immunology, whereas the complexity of cancer genesis requires an integrated analysis of heterogeneous data from several sources. Here we present the cancer-associated protein database (CAP), a novel analysis system for cancer-related data. CAP integrates data from multiple external databases, augments these data with functional annotations, and offers tools for statistical analysis of these data. We have employed CAP to analyze genes that have been found to cause an autoimmune response in cancer. In particular, we explored the connection between the autoimmune response, mutations, and overexpression of these genes. Our preliminary results suggest that mutations are not significant contributors to raising an antibody response against tumor antigens, whereas overexpression seems to play a more important role. We hereby demonstrate how different types of data can be integrated and analyzed successfully, providing interesting results. As the amount of available data is growing rapidly, a combined analysis will play an important role in exploring the genetic and immunological basis of cancer. CAP is freely available at the following web site: http://www.bioinf.uni-sb.de/CAP/.

Autoimmunity↗

Genetic Contributors to Postoperative Delirium and Their Implications for Dementia Outcomes.

BACKGROUND: Postoperative delirium (POD) is a perioperative neurocognitive disorder that substantially impairs patient recovery. Unfortunately, its genetic risk profile and relationship with subsequent dementia remain unclear. This study aimed to elucidate genetic contributors to POD identified via Hospital Episode Statistics codes and to examine its association with subsequent dementia. METHODS: The study included 230,179 noncardiac and 21,254 cardiac surgery subjects from the UK Biobank, defining POD using delirium codes from the International Classification of Diseases (10th revision) recorded within the first 7 postoperative days. Genome-wide association studies were performed in the noncardiac and cardiac cohorts and their prespecified subgroups, followed by functional annotation, gene prioritization and drug-target analyses. Associations between POD and subsequent dementia were estimated using Cox models. RESULTS: In the noncardiac cohort, one genome-wide significant locus was identified at the APOE region, with rs429358 as the lead variant ( P = 5.00&#x2009;&#xd7;&#x2009;10 -28 ). Integrative gene prioritization analyses highlighted multiple genes within this locus. Exploratory drug-target analyses suggested potential subgroup-specific drug-target enrichment. In the cardiac cohort, no genome-wide significant signals were detected. POD was associated with all-cause dementia after both noncardiac (hazard ratio, 6.45; 95% CI, 5.45 to 7.63) and cardiac (hazard ratio, 2.95; 95% CI, 1.71 to 5.08) surgeries. CONCLUSIONS: This study demonstrates APOE as a genetic risk locus for International Classification of Diseases-coded POD in the noncardiac surgery setting and confirms an association between POD and subsequent dementia.

Humans↗

Common genetic mechanisms between obesity and COVID-19 severity: unravelling pleiotropic loci and biological pathways.

COVID-19 and obesity are complex conditions marked by immune and metabolic dysfunction, with the former still ranking among the leading causes of death from infectious diseases worldwide and the latter reaching pandemic proportions. Clinical evidence consistently shows that obesity increases the risk of severe COVID-19, yet the biological mechanisms underlying this association remain unclear. Given their physiological and clinical overlap, they may share genetic pathways. We investigated genetic variants jointly associated with body mass index (BMI) and COVID-19 using publicly available genome-wide data. A conjunctional false discovery rate (conjFDR) approach identified shared variants between BMI and three COVID-19 phenotypes: infection, hospitalization and very severe respiratory illness. Functional annotation and pathway enrichment analyses were performed to explore the biological context of these variants, followed by a phenome-wide association study (PheWAS) to characterize pleiotropy. Shared variants were enriched in immune, metabolic and hormonal signaling pathways, including metal ion transport and glycosylation. The overlap with BMI was strongest for hospitalized and severe cases, suggesting common mechanisms underlying disease progression rather than infection. These findings suggest a biologically meaningful genetic overlap between obesity and COVID-19 severity, highlighting pleiotropy as a key feature in complex disease interactions and potential shared therapeutic targets.

BMI↗

Quantitative evolutionary genomics: differential gene expression and male reproductive success in Drosophila melanogaster.

We combined traditional quantitative genetics and oligonucleotide microarrays to examine within-population genetic variation in a trait closely related to fitness. The trait, male reproductive success under competitive conditions (MCRS), is of central importance to both life-history and sexual-selection theory. We identified 27 candidate genes whose expression levels were associated with within-population variation in MCRS. "High" MCRS was associated with low expression of a cytochrome P450 that causes pesticide resistance, suggesting a fitness cost to resistance. Two groups of metabolic proteins (glutathione transferases and phosphatases) were significantly over-represented, and a large portion of the candidates are genes involved in oxidative stress resistance, energy acquisition or energy storage. Genes expressed in accessory glands and testes were not over-represented among differentially expressed genes, but testis-expressed genes were significantly more likely to be upregulated in high MCRS genotypes. Finally, nine candidate genes that we identified had no previous functional annotation, and this experiment suggests that they play a role in male reproductive success.

Animals↗

Molecular control of the oocyte to embryo transition.

The elucidation of the molecular control of the initiation of mammalian embryogenesis is possible now that the transcriptomes of the full-grown oocyte and two-cell stage embryo have been prepared and analysed. Functional annotation of the transcriptomes using gene ontology vocabularies, allows comparison of the oocyte and two-cell stage embryo between themselves, and with all known mouse genes in the Mouse Genome Database. Using this methodology one can outline the general distinguishing features of the oocyte and the two-cell stage embryo. This, when combined with oocyte-specific targeted deletion of genes, allows us to dissect the molecular networks at play as the differentiated oocyte and sperm transit into blastomeres with unlimited developmental potential.

Animals↗

Microbiology Galaxy Lab: The first community-driven gateway for reproducible and FAIR analysis of microbial data.

The explosion of microbial omics data has outpaced the ability of many researchers to analyze it, with complex tools and limited computational resources creating barriers to discovery. To address this gap, we present the Microbiology Galaxy Lab: a free, globally accessible, community-supported platform that combines state-of-the-art analytical power with user-friendly accessibility. Supported by the Galaxy and global microbiology communities, this platform integrates over 315 tool suites and 115 curated workflows, enabling comprehensive metabarcoding, (meta)genomic, (meta)transcriptomic, and (meta)proteomic data analysis within a FAIR-aligned environment. It also supports research in the health and infectious disease sectors, as well as in environmental microbiology. The platform's utility is exemplified through various use cases, including antimicrobial resistance tracking, biomarker prediction, microbiome classification, and functional annotation of key microbes. Built on reproducibility and community engagement, it supports creation, sharing, and updating of best-practice workflows. Over 35 tutorials and learning paths empower scientists, fostering an ecosystem that keeps resources at the forefront of microbial science. The Microbiology Galaxy Lab enables collective analysis, democratising research, thereby accelerating discovery across the global microbiology community (microbiology.usegalaxy.org, .eu, .org.au, .fr).

Journal Article↗

An allelic resolution gene atlas for tetraploid potato provides insights into tuberization and stress resilience.

Tubers are modified underground stems that enable asexual, clonal reproduction and serve as a mechanism for overwintering and avoidance of herbivory. Tubers are wide-spread across angiosperms with some species such as Solanum tuberosum L. (potato) serving as a vital crop for human consumption. Genes responsible for tuber initiation and disease resistance have been characterized in potato including StSP6A, a homolog of Flowering Time, that functions as tuberigen, the equivalent of florigen. To elucidate additional molecular and genetic mechanisms underlying potato biology including tuber initiation, tuber development, and stress responses, we generated a developmental and abiotic/biotic-stress gene expression atlas from 34 tissues and treatments of Atlantic, a tetraploid cultivar. Using the haplotype-phased tetraploid Atlantic genome assembly and expression abundances of 129,218 genes, we constructed gene coexpression modules that represent networks associated with distinct developmental stages as well as stress responses. Functional annotations were given to modules and used to identify genes involved in tuberization and stress resilience. Structural variation from a pan-genomic analysis across four cultivated potato genome assemblies as well as domestication and wild introgression data allowed for deeper insights into the modules to identify key genes involved in tuberization and stress responses. This study underscores the importance of transcriptional regulation in tuberization and provides a comprehensive framework for future research on potato development and improvement.

Journal Article↗