Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,783 records · Page 99Linked to original sources

A phylogenomic gene cluster resource: the Phylogenetically Inferred Groups (PhIGs) database.

BACKGROUND: We present here the PhIGs database, a phylogenomic resource for sequenced genomes. Although many methods exist for clustering gene families, very few attempt to create truly orthologous clusters sharing descent from a single ancestral gene across a range of evolutionary depths. Although these non-phylogenetic gene family clusters have been used broadly for gene annotation, errors are known to be introduced by the artifactual association of slowly evolving paralogs and lack of annotation for those more rapidly evolving. A full phylogenetic framework is necessary for accurate inference of function and for many studies that address pattern and mechanism of the evolution of the genome. The automated generation of evolutionary gene clusters, creation of gene trees, determination of orthology and paralogy relationships, and the correlation of this information with gene annotations, expression information, and genomic context is an important resource to the scientific community. DISCUSSION: The PhIGs database currently contains 23 completely sequenced genomes of fungi and metazoans, containing 409,653 genes that have been grouped into 42,645 gene clusters. Each gene cluster is built such that the gene sequence distances are consistent with the known organismal relationships and in so doing, maximizing the likelihood for the clusters to represent truly orthologous genes. The PhIGs website contains tools that allow the study of genes within their phylogenetic framework through keyword searches on annotations, such as GO and InterPro assignments, and sequence similarity searches by BLAST and HMM. In addition to displaying the evolutionary relationships of the genes in each cluster, the website also allows users to view the relative physical positions of homologous genes in specified sets of genomes. SUMMARY: Accurate analyses of genes and genomes can only be done within their full phylogenetic context. The PhIGs database and corresponding website http://phigs.org address this problem for the scientific community. Our goal is to expand the content as more genomes are sequenced and use this framework to incorporate more analyses.

Base Sequence↗

Functional insights from the molecular modelling of a novel two-component system.

Two-component systems (TCSs) are the major signalling pathway in bacteria and represent potential drug targets. Among the 11 paired TCS proteins present in Mycobacterium tuberculosis H37Rv, the histidine kinases (HKs) Rv0600c (HK1) and Rv0601c (HK2) are annotated to phosphorylate one response regulator (RR) Rv0602c (TcrA). We wanted to establish the sequence-structure-function relationship to elucidate the mechanism of phosphotransfer using in silico methods. Sequence alignments and codon usage analysis showed that the two domains encoded by a single gene in homologous HKs have been separated into individual open-reading frames in M. tuberculosis. This is the first example where two incomplete HKs are involved in phosphorylating a single RR. The model shows that HK2 is a unique histidine phosphotransfer (HPt)-mono-domain protein, not found as lone protein in other bacteria. The secondary structure of HKs was confirmed using "far-UV" circular dichroism study of purified proteins. We propose that HK1 phosphorylates HK2 at the conserved H131 and the phosphoryl group is then transferred to D73 of TcrA.

Amino Acid Sequence↗

Plant receptor-like kinase gene family: diversity, function, and signaling.

Plant receptor-like kinases (RLKs) are transmembrane proteins with putative amino-terminal extracellular domains and carboxyl-terminal intracellular kinase domains, with striking resemblance in domain organization to the animal receptor tyrosine kinases such as epidermal growth factor receptor. The recently sequenced Arabidopsis genome contains more than 600 RLK homologs, representing nearly 2.5% of the annotated protein-coding genes in Arabidopsis. Although only a handful of these genes have known functions and fewer still have identified ligands or downstream targets, the studies of several RLKs such as CLAVATA1, Brassinosteroid Insensitive 1, Flagellin Insensitive 2, and S-locus receptor kinase provide much-needed information on the functions mediated by members of this large gene family. RLKs control a wide range of processes, including development, disease resistance, hormone perception, and self-incompatibility. Combined with the expression studies and biochemical analysis of other RLKs, more details of RLK function and signaling are emerging.

Animals↗

Onto-Tools: an ensemble of web-accessible, ontology-based tools for the functional design and interpretation of high-throughput gene expression experiments.

The Onto-Tools suite is composed of an annotation database and five seamlessly integrated web-accessible data mining tools: Onto-Express (OE), Onto-Compare (OC), Onto-Design (OD), Onto-Translate (OT) and Onto-Miner (OM). OM is a new tool that provides a unified access point and an application programming interface for most annotations available. Our database has been enhanced with more than 120 new commercial microarrays and annotations for Rattus norvegicus, Drosophila melanogaster and Carnorhabditis elegans. The Onto-Tools have been redesigned to provide better biological insight, improved performance and user convenience. The new features implemented in OE include support for gene names, LocusLink IDs and Gene Ontology (GO) IDs, ability to specify fold changes for the input genes, links to the KEGG pathway database and detailed output files. OC allows comparisons of the functional bias of more than 170 commercial microarrays. The latest version of OD allows the user to specify keywords if the exact GO term is not known as well as providing more details than the previous version. OE, OC and OD now have an integrated GO browser that allows the user to customize the level of abstraction for each GO category. The Onto-Tools are available online at http://vortex.cs.wayne.edu/Projects.html.

Animals↗

Ureaplasma urealyticum: an opportunity for combinatorial genomics.

Examination of genomic or enzymatic activity data alone neither provides a complete picture of metabolic function or potential nor confidently reveals sites amenable to inhibition. Furthermore, in some cases, gene annotation and in aqua assays disagree by describing gene annotation without enzyme activity and enzyme activity without homologous annotation. The newly sequenced genome of Ureaplasma urealyticum (parvum) is another prokaryote example of the class Mollicutes where such confounding differences are observed. The little-considered role of some proteins as multifunctional enzymes - substitutes for 'missing' genes - could both partially explain the apparent anomalies and relate to any inaccurate deductions of inhibitor function. A combinatorial analysis involving available evidence of genomic sequence, transcription, translational phenomena, structure and enzymatic activity gives the best picture of the organism's vital metabolic alternatives.

Bacterial Proteins↗

SPD--a web-based secreted protein database.

With the improved secreted protein prediction approach and comprehensive data sources, including Swiss-Prot, TrEMBL, RefSeq, Ensembl and CBI-Gene, we have constructed secretomes of human, mouse and rat, with a total of 18 152 secreted proteins. All the entries are ranked according to the prediction confidence. They were further annotated via a proteome annotation pipeline that we developed. We also set up a secreted protein classification pipeline and classified our predicted secreted proteins into different functional categories. To make the dataset more convincing and comprehensive, nine reference datasets are also integrated, such as the secreted proteins from the Gene Ontology Annotation (GOA) system at the European Bioinformatics Institute, and the vertebrate secreted proteins from Swiss-Prot. All these entries were grouped via a TribeMCL based clustering pipeline. We have constructed a web-based secreted protein database, which has been publicly available at http://spd.cbi.pku.edu.cn. Users can browse the database via a GO assignment or chromosomal-location-based interface. Moreover, text query and sequence similarity search are also provided, and the sequence and annotation data can be downloaded freely from the SPD website.

Animals↗

TET2 promotes monocyte inflammatory activation in asthma via ALKBH5-m6A regulation and PI3K signaling: evidence from m6A-SNP and single-cell analyses.

Asthma is a complex inflammatory airway disease with strong genetic determinants, yet the functional relevance of most asthma-associated non-coding variants remains unclear. Emerging evidence suggests that N6-methyladenosine (m6A) modification may serve as a critical epitranscriptomic link between genetic variation and immune regulation. In this study, we aimed to systematically identify functionally relevant m6A-regulated genes in asthma by integrating large-scale GWAS data, m6A-SNP annotations, and single-cell transcriptomic analyses, and to investigate their roles in monocyte-driven airway inflammation. We identified TET2 as a key m6A-regulated gene associated with both asthma and lung function, which was selectively upregulated in monocytes during asthma and accompanied by activation of inflammatory and PI3K signaling pathways. Mechanistic experiments further demonstrated that inflammatory stimulation induced ALKBH5 expression, reduced m6A modification of TET2 mRNA, and increased TET2 protein levels, thereby promoting PI3K/AKT signaling and pro-inflammatory cytokine production, whereas inhibition of TET2 or ALKBH5 attenuated these effects. Collectively, these findings demonstrate that ALKBH5-mediated m6A regulation of TET2 enhances PI3K/AKT signaling in monocytes, thereby promoting inflammatory responses in asthma. Our study establishes TET2 as a key m6A-regulated gene linking genetic susceptibility to monocyte-driven inflammation, and highlights the ALKBH5-m6A-TET2 axis as a potential therapeutic target for modulating aberrant immune responses in asthma.

Humans↗

The first archaeal ATP-dependent glucokinase, from the hyperthermophilic crenarchaeon Aeropyrum pernix, represents a monomeric, extremely thermophilic ROK glucokinase with broad hexose specificity.

An ATP-dependent glucokinase of the hyperthermophilic aerobic crenarchaeon Aeropyrum pernix was purified 230-fold to homogeneity. The enzyme is a monomeric protein with an apparent molecular mass of about 36 kDa. The apparent K(m) values for ATP and glucose (at 90 degrees C and pH 6.2) were 0.42 and 0.044 mM, respectively; the apparent V(max) was about 35 U/mg. The enzyme was specific for ATP as a phosphoryl donor, but showed a broad spectrum for phosphoryl acceptors: in addition to glucose, which showed the highest catalytic efficiency (k(cat)/K(m)), the enzyme also phosphorylates glucosamin, fructose, mannose, and 2-deoxyglucose. Divalent cations were required for maximal activity: Mg(2+), which was most effective, could partially be replaced with Co(2+), Mn(2+), and Ni(2+). The enzyme had a temperature optimum of at least 100 degrees C and showed significant thermostability up to 100 degrees C. The coding function of open reading frame (ORF) APE2091 (Y. Kawarabayasi, Y. Hino, H. Horikawa, S. Yamazaki, Y. Haikawa, K. Jin-no, M. Takahashi, M. Sekine, S. Baba, A. Ankai, H. Kosugi, A. Hosoyama, S. Fukui, Y. Nagai, K. Nishijima, H. Nakazawa, M. Takamiya, S. Masuda, T. Funahashi, T. Tanaka, Y. Kudoh, J. Yamazaki, N. Kushida, A. Oguchi, and H. Kikuchi, DNA Res. 6:83-101, 145-152, 1999), previously annotated as gene glk, coding for ATP-glucokinase of A. pernix, was proved by functional expression in Escherichia coli. The purified recombinant ATP-dependent glucokinase showed a 5-kDa higher molecular mass on sodium dodecyl sulfate-polyacrylamide gel electrophoresis, but almost identical kinetic and thermostability properties in comparison to the native enzyme purified from A. pernix. N-terminal amino acid sequence of the native enzyme revealed that the translation start codon is a GTG 171 bp downstream of the annotated start codon of ORF APE2091. The amino acid sequence deduced from the truncated ORF APE2091 revealed sequence similarity to members of the ROK family, which comprise bacterial sugar kinases and transcriptional repressors. This is the first report of the characterization of an ATP-dependent glucokinase from the domain of Archaea, which differs from its bacterial counterparts by its monomeric structure and its broad specificity for hexoses.

Adenosine Triphosphate↗

MJ0109 is an enzyme that is both an inositol monophosphatase and the 'missing' archaeal fructose-1,6-bisphosphatase.

In sequenced genomes, protein coding regions with unassigned function constitute between 10 and 50% of all open reading frames. Often key enzymes cannot be identified using sequence homology searches. For example, despite the fact that methanogens have an apparently functional gluconeogenesis pathway, standard tools have been unable to identify a fructose-1,6-bisphosphatase (FBPase) gene in the sequenced Methanoccocus jannaschii genome. Using a combination of functional and structural tools, we have shown that the protein product of the M. jannaschii gene MJ0109, which had been tentatively annotated as an inositol monophosphatase (IMPase), has both IMPase and FBPase activities. Moreover, several gene products annotated as IMPases from different thermophilic organisms also possess FBPase activity. Thus, we have found the FBPase that was 'missing' in thermophiles and shown that it also functions as an IMPase.

Amino Acid Sequence↗

Human promoter genomic composition demonstrates non-random groupings that reflect general cellular function.

BACKGROUND: The purpose of this study is to determine whether or not there exists nonrandom grouping of cis-regulatory elements within gene promoters that can be perceived independent of gene expression data and whether or not there is any correlation between this grouping and the biological function of the gene. RESULTS: Using ProSpector, a web-based promoter search and annotation tool, we have applied an unbiased approach to analyze the transcription factor binding site frequencies of 1400 base pair genomic segments positioned at 1200 base pairs upstream and 200 base pairs downstream of the transcriptional start site of 7298 commonly studied human genes. Partitional clustering of the transcription factor binding site composition within these promoter segments reveals a small number of gene groups that are selectively enriched for gene ontology terms consistent with distinct aspects of cellular function. Significance ranking of the class-determining transcription factor binding sites within these clusters show substantial overlap between the gene ontology terms of the transcriptions factors associated with the binding sites and the gene ontology terms of the regulated genes within each group. CONCLUSION: Thus, gene sorting by promoter composition alone produces partitions in which the "regulated" and the "regulators" cosegregate into similar functional classes. These findings demonstrate that the transcription factor binding site composition is non-randomly distributed between gene promoters in a manner that reflects and partially defines general gene class function.

Binding Sites↗

Genome-wide screening and functional analysis of protein glycosylation-related genes involved in tomato fruit ripening.

Protein glycosylation, an essential co- and post-translational modification, plays critical roles in plant growth, development, and stress responses. However, its functional role in tomato fruit ripening has not been extensively investigated. Here, key protein glycosylation-related genes involved in tomato fruit ripening were identified by genome-wide screen and subsequently functional characterization. First, a dataset comprising 242 glycosylation-related proteins was established based on Gene Ontology annotations in tomato, combined with sequence homology to protein glycosylation-related proteins from Arabidopsis thaliana and Homo sapiens. Then, Subsequently, 28 genes encoding highly expressed glycosylation-related proteins (RPKM > 30) at the breaker (BR) stage were selected for functional screening, and subsequently 6 genes were identified as regulators of fruit ripening by method of virus-induced gene silencing (VIGS). Among them, Solyc03g098600 (STT3B), Solyc01g109410 (OST48), Solyc04g082670 (RPN1), and Solyc08g076460 (DAD1) functioned as positive regulators of tomato fruit ripening, whereas Solyc04g005340 (UAM2) and Solyc08g075340 (XEG113), acted as negative regulators. The expression of these genes responded dynamically to multiple ripening-related cues, including temperature, light, ethylene, and transcription factors. Furthermore, silencing of these genes individually affected the expression of genes involved in fruit ripening, including ethylene biosynthesis genes (ACS2, ACS4, ACO1, and ACO3), ripening-associated transcription factors (RIN, NOR, NOR-LIKE1, FUL1, and FUL2), and the key gene (PSY1) of lycopene biosynthesis pathway. Collectively, these findings demonstrate that protein glycosylation plays an important role in tomato fruit ripening by modulating ethylene signaling, ripening-associated transcriptional regulation, and lycopene biosynthesis.

Fruit ripening↗

Orthology between the genomes of Plasmodium falciparum and rodent malaria parasites: possible practical applications.

The work of the consortium that has been formed to complete the entire sequence of the genome of a selected clone of the human malaria parasite, Plasmodium falciparum, is almost finished. Already huge tracts of the genome are available as fully assembled chromosomes or large contigs and the work of initial annotation is in an advanced state. Post-genomic research is in one sense the process of furthering the process of annotation, creating biological atlases and preliminary attempts to make global descriptions of gene transcription and proteome analysis are underway. Comparison between significant amounts of genome data from both closely, and more distantly related organisms, can facilitate the identification of genes themselves, coordinately regulated gene expression groups, gene function and genome organization. Models of malaria can fulfil these functions and in addition provide an experimental system wherein predictions can be tested and basic experimental investigations performed within numerous aspects of disease, pathology, parasite-host and parasite-vector interactions. Comparative genomics in Plasmodium has already been shown to have informative roles in the completion of annotation and the elucidation of gene function. These roles will be illustrated by example and used as the basis for a discussion of the utility of genome information and malaria models in realizing the desired product of Plasmodium genomics, the development of malaria therapies.

Animals↗

Persistent biases in the amino acid composition of prokaryotic proteins.

Correspondence analysis of 28 proteomes selected to span the entire realm of prokaryotes revealed universal biases in the proteins' amino acid distribution. Integral Inner Membrane Proteins always form an individual cluster, which can then be used to predict protein localisation in unknown proteomes, independently of the organism's biotope or kingdom. Orphan proteins are consistently rich in aromatic residues. Another bias is also ubiquitous: the amino acid composition is driven by the G + C content of the first codon position. An unexpected bias is driven, in many proteomes, by the AAN box of the genetic code, suggesting some functional biochemical relationship between asparagine and lysine. Less-significant biases are driven by the rare amino acids, cysteine and tryptophan. Some allow identification of species-specific functions or localisation such as surface or exported proteins. Errors in genome annotations are also revealed by correspondence analysis, making it useful for quality control and correction.

Amino Acids↗

Targeted overexpression of the Escherichia coli MinC protein in higher plants results in abnormal chloroplasts.

Higher plant chloroplast division involves some of the same types of proteins that are required in prokaryotic cell division. These include two of the three Min proteins, MinD and MinE, encoded by the min operon in bacteria. Noticeably absent from annotated sequences from higher plants is a MinC homologue. A higher plant functional MinC homologue that would interfere with FtsZ polymerization, has yet to be identified. We sought to determine whether expression of the bacterial MinC in higher plants could affect chloroplast division. The Escherichia coli minC (EcMinC) gene was isolated and inserted behind the Arabidopsis thaliana RbcS transit peptide sequence for chloroplast targeting. This TP-EcMinC gene driven by the CaMV 35S(2) constitutive promoter was then transformed into tobacco (Nicotiana tabacum L.). Abnormally large chloroplasts were observed in the transgenic plants suggesting that overexpression of the E. coli MinC perturbed higher plant chloroplast division.

Chloroplasts↗

Engineered Lactiplantibacillus plantarum and Levilactobacillus brevis utilizing ribonucleoprotein-mediated editing for inactivation of hemolysin gene.

Lactiplantibacillus plantarum and Levilactobacillus brevis are widely used probiotics with significant potential as chassis organisms for probiotic engineering. However, their bioengineering remains underdeveloped compared to that of other probiotic bacteria due to the limited availability of genetic tools. Although CRISPR-Cas systems have shown promise for genome editing in Lactobacillus species, strain- or site-specific targeting challenges must be overcome to enhance their broader applicability. This study aimed to develop a novel editing system with reduced dependency on plasmids and antibiotics in L. plantarum WCFS1, L. plantarum SPC 72 - 1 and L. brevis SPC-SNU 70 - 2 using a Cas9-gRNA ribonucleoprotein (RNP) complex. Although the hlyIII gene has been annotated as a hemolysin-related gene in several Lactobacillus genomes, no functional hemolytic activity has been definitively demonstrated to date. In this study, hlyIII was selected as a target to evaluate genome editing efficiency and to assess its potential relevance to strain safety. To construct ΔhlyIII strains, the RNP complex targeting hlyIII was separately transformed with recombinase RecE/T and double-stranded donor DNA. As a result, ΔhlyIII mutants were obtained under optimized electroporation conditions. Sequencing analysis revealed a 50 bp deletion and the introduction of a stop codon in hlyIII across all mutant strains. The hemolytic activity test showed a reduction in free hemoglobin levels in the ΔhlyIII strains compared to the wild type: 27.0%, 74.3%, and 5.0% in L. plantarum WCFS1, L. plantarum SPC 72 - 1, and L. brevis SPC-SNU 70 - 2, respectively. These results suggest strain-dependent differences in hemolytic activity and indicate that inactivation of hlyIII may contribute to reduced hemolysis, although further validation is needed to clarify its functional role. In conclusion, the hlyIII gene was successfully edited in L. plantarum and L. brevis using Cas9-gRNA ribonucleoprotein-mediated editing, demonstrating the feasibility of this genome editing platform for application in probiotic strains.

Gene Editing↗

Identification and analysis of the mouse basic/Helix-Loop-Helix transcription factor family.

The basic/Helix-Loop-Helix (bHLH) proteins are a family of transcription factors that regulates a variety of biological processes. Based on a previously defined consensus motif, we identified the complete set of bHLH protein family from the mouse proteome databases and carried out a series of bioinformatics analysis. As results, 124 mouse bHLH proteins were identified in this study, and 28 of them were additional bHLH proteins beyond the previous report. These 124 mouse bHLH proteins were classified into groups from A to F by the nomenclature and phylogenetic analysis. Statistic analysis of the Gene Ontology annotation of these proteins showed that the bHLH proteins tend to perform functions related to cell differentiation and development. Gene function enrichment analysis among six groups illuminated that the proteins in certain group tend to have special biology functions, so that the molecular function of the uncharacterized proteins in groups could be inferred.

Amino Acid Sequence↗

The genomic sequence and analysis of the swine major histocompatibility complex.

We describe the generation and analysis of an integrated sequence map of a 2.4-Mb region of pig chromosome 7, comprising the classical class I region, the extended and classical class II regions, and the class III region of the major histocompatibility complex (MHC), also known as swine leukocyte antigen (SLA) complex. We have identified and manually annotated 151 loci, of which 121 are known genes (predicted to be functional), 18 are pseudogenes, 8 are novel CDS loci, 3 are novel transcripts, and 1 is a putative gene. Nearly all of these loci have homologues in other mammalian genomes but orthologues could be identified with confidence for only 123 genes. The 28 genes (including all the SLA class I genes) for which unambiguous orthology to genes within the human reference MHC could not be established are of particular interest with respect to porcine-specific MHC function and evolution. We have compared the porcine MHC to other mammalian MHC regions and identified the differences between them. In comparison to the human MHC, the main differences include the absence of HLA-A and other class I-like loci, the absence of HLA-DP-like loci, and the separation of the extended and classical class II regions from the rest of the MHC by insertion of the centromere. We show that the centromere insertion has occurred within a cluster of BTNL genes located at the boundary of the class II and III regions, which might have resulted in the loss of an orthologue to human C6orf10 from this region.

Animals↗

Trait-to-gene: a computational method for predicting the function of uncharacterized genes.

The function of unknown genes is often inferred from comparisons to well-characterized homologs. In this paper, we show that, even if all of the homologs of a gene are unannotated, its function may be deduced through phylogenetic profiling. We have designed a series of algorithms that make functional predictions of genes based on orthology and set theory, but our approach to predicting gene function requires no previous knowledge of homolog function. With this technique, we successfully identified 94% of the clusters of orthologous groups that are known to be involved in flagella development or function. As a test, we removed the function of three putative flagellar genes that had been previously uncharacterized in Bacillus subtilis. We observed a motility phenotype for two of these three genes. Thus, these algorithms allow for high-throughput functional prediction of genes beyond that provided by simple orthology-based annotation endeavors.

Algorithms↗