Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Chromosome-level genome assembly of Cheilinus chlorourus (Bloch, 1791) (Perciformes: Labridae).

In the classification of marine fish, the Labridae family ranks second in terms of species diversity and plays a vital role in coral reef ecosystems, comprising over 600 species across 82 genera. Despite its significance for ecological and evolutionary studies, genomic research on this group has lagged, resulting in a shortage of data, particularly regarding high-quality chromosome-level genome assemblies. To address this gap, this study focused on Cheilinus chlorourus from the Labridae family and successfully achieved a chromosome-level genome assembly. By integrating Illumina, PacBio, and Hi-C sequencing data, we assembled a genome measuring 940.36 Mb, with 926.86 Mb (98.56%) of the gene assembly organized into 21 chromosomes. A total of 29,213 protein-coding genes (PCGs) were identified, and 79.93% of these genes were functionally annotated. With this high-quality genome assembly, future investigations into the functional genomics and ecology of C. chlorourus will have a solid scientific foundation.

Animals↗

ProtoNet: hierarchical classification of the protein space.

The ProtoNet site provides an automatic hierarchical clustering of the SWISS-PROT protein database. The clustering is based on an all-against-all BLAST similarity search. The similarities' E-score is used to perform a continuous bottom-up clustering process by applying alternative rules for merging clusters. The outcome of this clustering process is a classification of the input proteins into a hierarchy of clusters of varying degrees of granularity. ProtoNet (version 1.3) is accessible in the form of an interactive web site at http://www.protonet.cs.huji.ac.il. ProtoNet provides navigation tools for monitoring the clustering process with a vertical and horizontal view. Each cluster at any level of the hierarchy is assigned with a statistical index, indicating the level of purity based on biological keywords such as those provided by SWISS-PROT and InterPro. ProtoNet can be used for function prediction, for defining superfamilies and subfamilies and for large-scale protein annotation purposes.

Animals↗

Identification of a novel gene, URE2, that functionally complements a urease-negative clinical strain of Cryptococcus neoformans.

A urease-negative serotype A strain of Cryptococcus neoformans (B-4587) was isolated from the cerebrospinal fluid of an immunocompetent patient with a central nervous system infection. The URE1 gene encoding urease failed to complement the mutant phenotype. Urease-positive clones of B-4587 obtained by complementing with a genomic library of strain H99 harboured an episomal plasmid containing DNA inserts with homology to the sudA gene of Aspergillus nidulans. The gene harboured by these plasmids was named URE2 since it enabled the transformants to grow on media containing urea as the sole nitrogen source while the transformants with an empty vector failed to grow. Transformation of strain B-4587 with a plasmid construct containing a truncated version of the URE2 gene failed to complement the urease-negative phenotype. Disruption of the native URE2 gene in a wild-type serotype A strain H99 and a serotype D strain LP1 of C. neoformans resulted in the inability of the strains to grow on media containing urea as the sole nitrogen source, suggesting that the URE2 gene product is involved in the utilization of urea by the organism. Virulence in mice of the urease-negative isolate B-4587, the urease-positive transformants containing the wild-type copy of the URE2 gene, and the urease-negative vector-only transformants was comparable to that of the H99 strain of C. neoformans regardless of the infection route. Virulence of the URE2 disruption stain of H99 was slightly reduced compared to the wild-type strain in the intravenous model but was significantly attenuated in the inhalation model. These results indicate that the importance of urease activity in pathogenicity varies depending on the strains of C. neoformans used and/or the route of infection. Furthermore, this study shows that complementation cloning can serve as a useful tool to functionally identify genes such as URE2 that have otherwise been annotated as hypothetical proteins in genomic databases.

Animals↗

Genome sequence of Haloarcula marismortui: a halophilic archaeon from the Dead Sea.

We report the complete sequence of the 4,274,642-bp genome of Haloarcula marismortui, a halophilic archaeal isolate from the Dead Sea. The genome is organized into nine circular replicons of varying G+C compositions ranging from 54% to 62%. Comparison of the genome architectures of Halobacterium sp. NRC-1 and H. marismortui suggests a common ancestor for the two organisms and a genome of significantly reduced size in the former. Both of these halophilic archaea use the same strategy of high surface negative charge of folded proteins as means to circumvent the salting-out phenomenon in a hypersaline cytoplasm. A multitiered annotation approach, including primary sequence similarities, protein family signatures, structure prediction, and a protein function association network, has assigned putative functions for at least 58% of the 4242 predicted proteins, a far larger number than is usually achieved in most newly sequenced microorganisms. Among these assigned functions were genes encoding six opsins, 19 MCP and/or HAMP domain signal transducers, and an unusually large number of environmental response regulators-nearly five times as many as those encoded in Halobacterium sp. NRC-1--suggesting H. marismortui is significantly more physiologically capable of exploiting diverse environments. In comparing the physiologies of the two halophilic archaea, in addition to the expected extensive similarity, we discovered several differences in their metabolic strategies and physiological responses such as distinct pathways for arginine breakdown in each halophile. Finally, as expected from the larger genome, H. marismortui encodes many more functions and seems to have fewer nutritional requirements for survival than does Halobacterium sp. NRC-1.

Archaeal Proteins↗

Characterization of a novel Neisseria meningitidis Fur and iron-regulated operon required for protection from oxidative stress: utility of DNA microarray in the assignment of the biological role of hypothetical genes.

We have previously shown that in the human pathogen Neisseria meningitidis group B (MenB) more than 200 genes are regulated in response to growth with iron. Among the Fur-dependent, upregulated genes identified by microarray analysis was a putative operon constituted by three genes, annotated as NMB1436, NMB1437 and NMB1438 and encoding proteins with so far unknown function. The operon was remarkably upregulated in the presence of iron and, on the basis of gel retardation analysis, its regulation was Fur dependent. In this study, we have further characterized the role of iron and Fur in the regulation of the NMB1436-38 operon and we have mapped the promoter and the Fur binding site. We also demonstrate by mutant analysis that the NMB1436-38 operon is required for protection of MenB to hydrogen peroxide-mediated killing. By using both microarray analysis and S1 mapping, we demonstrate that the operon is not regulated by oxidative stress signals. We also show that the deletion of the NMB1436-38 operon results in an impaired capacity of MenB to survive in the blood of mice using an adult mouse model of MenB infection. Finally, we show that the NMB1436-38 deletion mutant exhibits increased susceptibility to the killing activity of polymorphonuclears (PMNs), suggesting that the 'attenuated' phenotype is mediated in part by the increased sensitivity to reactive oxygen species-producing cells. This study represents one of the first examples of the use of DNA microarray to assign a biological role to hypothetical genes in bacteria.

Animals↗

Quantitative proteomic analysis of the brain reveals the potential antidepressant mechanism of Jiawei Danzhi Xiaoyao San in a chronic unpredictable mild stress mouse model of depression.

OBJECTIVE: To reveal the antidepressant mechanisms of Jiawei DanZhiXiaoYaoSan (,JD) in chronic unpredictable mild stress (CUMS)-induced depression in mice. METHODS: Using the CUMS mouse model of depression, the antidepressant effects of JD were assessed using the sucrose preference test (SPT), forced swimming test (FST), and tail suspension test (TST). Tandem mass tag (TMT)-based quantitative proteomic analysis of the brain was performed following JD treatment. Hierarchical clustering, Gene Ontology function annotation, Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment, and protein-protein interactions (PPIs) were used to analyze differentially expressed proteins (DEPs), which were further validated using quantitative real-time polymerase chain reaction (qRT-PCR) and Western blotting. RESULTS: Behavioral tests confirmed the anti-depressant effects of JD, and bioinformatics analysis revealed 59 DEPs, including 33 up-regulated and 26 down-regulated proteins, between the CUMS and JD-M groups. KEGG and PPI analyses revealed that neuro-filament proteins and the Ras signaling pathway may be key targets of JD in the treatment of depression. qRT-PCR and Western blotting results demonstrated that CUMS reduced the protein expression of neurofilament light (NEFL) and medium (NEFM) and inhibited the phosphorylation of extracellular regulated kinase 1/2 (ERK1/2), whereas JD promoted the phosphorylation of ERK1/2 and up-regulated the protein expression of NEFL and NEFM. CONCLUSIONS: The antidepressant mechanism of JD may be related to the up-regulation of p-ERK1/2 and neurofilament proteins.

Animals↗

Conserved codon composition of ribosomal protein coding genes in Escherichia coli, Mycobacterium tuberculosis and Saccharomyces cerevisiae: lessons from supervised machine learning in functional genomics.

Genomics projects have resulted in a flood of sequence data. Functional annotation currently relies almost exclusively on inter-species sequence comparison and is restricted in cases of limited data from related species and widely divergent sequences with no known homologs. Here, we demonstrate that codon composition, a fusion of codon usage bias and amino acid composition signals, can accurately discriminate, in the absence of sequence homology information, cytoplasmic ribosomal protein genes from all other genes of known function in Saccharomyces cerevisiae, Escherichia coli and Mycobacterium tuberculosis using an implementation of support vector machines, SVM(light). Analysis of these codon composition signals is instructive in determining features that confer individuality to ribosomal protein genes. Each of the sets of positively charged, negatively charged and small hydrophobic residues, as well as codon bias, contribute to their distinctive codon composition profile. The representation of all these signals is sensitively detected, combined and augmented by the SVMs to perform an accurate classification. Of special mention is an obvious outlier, yeast gene RPL22B, highly homologous to RPL22A but employing very different codon usage, perhaps indicating a non-ribosomal function. Finally, we propose that codon composition be used in combination with other attributes in gene/protein classification by supervised machine learning algorithms.

Algorithms↗

A Web-based classification system of DNA-binding protein families.

Rational classification of proteins encoded in sequenced genomes is critical for making the genome sequences maximally useful for functional and evolutionary studies. The family of DNA-binding proteins is one of the most populated and studied amongst the various genomes of bacteria, archaea and eukaryotes and the Web-based system presented here is an approach to their classification. The DnaProt resource is an annotated and searchable collection of protein sequences for the families of DNA-binding proteins. The database contains 3238 full-length sequences (retrieved from the SWISS-PROT database, release 38) that include, at least, a DNA-binding domain. Sequence entries are organized into families defined by PROSITE patterns, PRINTS motifs and de novo excised signatures. Combining global similarities and functional motifs into a single classification scheme, DNA-binding proteins are classified into 33 unique classes, which helps to reveal comprehensive family relationships. To maximize family information retrieval, DnaProt contains a collection of multiple alignments for each DNA-binding family while the recognized motifs can be used as diagnostically functional fingerprints. All available structural class representatives have been referenced. The resource was developed as a Web-based management system for online free access of customized data sets. Entries are fully hyperlinked to facilitate easy retrieval of the original records from the source databases while functional and phylogenetic annotation will be applied to newly sequenced genomes. The database is freely available for online search of a library containing specific patterns of the identified DNA-binding protein classes and retrieval of individual entries from our WWW server (http://kronos.biol.uoa.gr/~mariak/dbDNA.html).

Amino Acid Motifs↗

MobiDB-lite 4.0: faster prediction of intrinsic protein disorder and structural compactness.

MOTIVATION: In recent years, many disorder predictors have been developed to identify intrinsically disordered regions (IDRs) in proteins, achieving high accuracy. However, it may be difficult to interpret differences in predictions across methods. Consensus methods offer a simple solution, highlighting reliable predictions while filtering out uncertain positions. Here, we present a new version of MobiDB-lite, a consensus method designed to predict long IDRs and classify them based on compositional biases and conformational properties. RESULTS: MobiDB-lite 4.0 pipeline was optimized to be ten times faster than the previous version. It now provides compactness annotations based on predicted apparent scaling exponent. The newly added features and disorder subclassifications allow the users to get a comprehensive insight into the protein's function and characteristics. MobiDB-lite 4.0 is integrated into the MobiDB and DisProt databases. A version without the compactness predictor is integrated into InterProScan, propagating MobiDB-lite annotations to UniProtKB. AVAILABILITY AND IMPLEMENTATION: The MobiDB-lite 4.0 source code and a Docker container are available from the GitHub repository: https://github.com/BioComputingUP/MobiDB-lite.

Intrinsically Disordered Proteins↗

GeomeTRe: accurate calculation of geometrical descriptors of tandem repeat proteins.

MOTIVATION: Structured tandem repeat proteins (STRPs) are characterized by preserved structural motifs arranged in a modular way. The structural and functional diversity of STRPs makes them particularly important for studying evolution and novel structure-function relationships, and ultimately for designing new synthetic proteins with specific functions. One crucial aspect of their classification is the estimation of geometrical parameters, which can provide better insight into their properties and the relationship between the spatial arrangement of repeated units and protein function. Calculating geometric descriptors for STRPs is challenging because naturally occurring repeats are not "perfect" and often contain insertions and deletions. Existing tools for predicting structural symmetry work well on simple cases but often fail for most natural proteins. RESULTS: Here, we present GeomeTRe, an algorithm that calculates geometrical descriptors such as curvature (yaw), twist (roll), and pitch for a protein structure with known repeat unit positions. The algorithm simulates the movement of consecutive units, identifies rotational axes, and calculates the corresponding Tait-Bryan angles. GeomeTRe's parameters can enhance STRP annotation and classification by identifying variations in geometric arrangements among different functional groups. The package is fast and suitable for processing large protein structure datasets when repeat region information (e.g. from RepeatsDB) is available. AVAILABILITY AND IMPLEMENTATION: GeomeTRe is available as a Python package; source code and documentation can be found at https://github.com/BioComputingUP/GeomeTRe.

Algorithms↗

Protein coding potential of retroviruses and other transposable elements in vertebrate genomes.

We suggest an annotation strategy for genes encoded by retroviruses and transposable elements (RETRA genes) based on a set of marker protein domains. Usually RETRA genes are masked in vertebrate genomes prior to the application of automated gene prediction pipelines under the assumption that they provide no selective advantage to the host. Yet, we show that about 1000 genes in four vertebrate gene sets analyzed contain at least one RETRA gene marker domain. Using the conservation of genomic neighborhood (synteny), we were able to discriminate between RETRA genes with putative functionality in the vertebrates and those that probably function only in the context of mobile elements. We identified 35 such genes in human, along with their corresponding mouse and rat orthologs; which included almost all known human genes with similarity to mobile elements. The results also imply that the vast majority of the remaining RETRA genes in current gene sets are unlikely to encode vertebrate functions. To automatically annotate RETRA genes in other vertebrate genomes, we provide as a tool a set of marker protein domains and a manually refined list of domesticated or ancestral RETRA genes for rescuing genes with vertebrate functions.

Animals↗

Crystal structure of YecO from Haemophilus influenzae (HI0319) reveals a methyltransferase fold and a bound S-adenosylhomocysteine.

The crystal structure of YecO from Haemophilus influenzae (HI0319), a protein annotated in the sequence databases as hypothetical, and that has not been assigned a function, has been determined at 2.2-A resolution. The structure reveals a fold typical of S-adenosyl-L-methionine-dependent (AdoMet) methyltransferase enzymes. Moreover, a processed cofactor, S-adenosyl-L-homocysteine (AdoHcy), is bound to the enzyme, further confirming the biochemical function of HI0319 and its sequence family members. An active site arginine, shielded from bulk solvent, interacts with an anion, possibly a chloride ion, which in turn interacts with the sulfur atom of AdoHcy. The AdoHcy and nearby protein residues delineate a small solvent-excluded substrate binding cavity of 162 A(3) in volume. The environment surrounding the cavity indicates that the substrate molecule contains a hydrophobic moiety and an anionic group. Many of the residues that define the cavity are invariant in the HI0319 sequence family but are not conserved in other methyltransferases. Therefore, the substrate specificity of YecO enzymes is unique and differs from the substrate specificity of all other methyltransferases sequenced to date. Examination of the Enzyme Commission list of methyltransferases prompted a manual inspection of 10 possible substrates using computer graphics and suggested that the ortho-substituted benzoic acids fit best in the active site.

Binding Sites↗

The crystal structure of Rv1347c, a putative antibiotic resistance protein from Mycobacterium tuberculosis, reveals a GCN5-related fold and suggests an alternative function in siderophore biosynthesis.

Mycobacterium tuberculosis, the cause of tuberculosis, is a devastating human pathogen. The emergence of multidrug resistance in recent years has prompted a search for new drug targets and for a better understanding of mechanisms of resistance. Here we focus on the gene product of an open reading frame from M. tuberculosis, Rv1347c, which is annotated as a putative aminoglycoside N-acetyltransferase. The Rv1347c protein does not show this activity, however, and we show from its crystal structure, coupled with functional and bioinformatic data, that its most likely role is in the biosynthesis of mycobactin, the M. tuberculosis siderophore. The crystal structure of Rv1347c was determined by multiwavelength anomalous diffraction phasing from selenomethionine-substituted protein and refined at 2.2 angstrom resolution (r = 0.227, R(free) = 0.257). The protein is monomeric, with a fold that places it in the GCN5-related N-acetyltransferase (GNAT) family of acyltransferases. Features of the structure are an acyl-CoA binding site that is shared with other GNAT family members and an adjacent hydrophobic channel leading to the surface that could accommodate long-chain acyl groups. Modeling the postulated substrate, the N(epsilon)-hydroxylysine side chain of mycobactin, into the acceptor substrate binding groove identifies two residues at the active site, His130 and Asp168, that have putative roles in substrate binding and catalysis.

Acyl Coenzyme A↗

AI-enabled viral genomics: from virus discovery to host prediction and emerging variant forecasting.

The rapid expansion of metagenomic sequencing has generated vast repositories of viral sequence data that far outpace our capacity to interpret them using conventional approaches. Highly divergent sequences, sparse functional annotation, and taxonomically uneven sampling present fundamental challenges for reference-dependent methods, which lose sensitivity precisely for novel and understudied viruses with high public health relevance. Artificial intelligence (AI) provides a new avenue to address these challenges by enabling predictive inference from viral genomes and proteins while reducing dependence on sequence similarity. In this Review, we discuss representative advances in AI for virus discovery, taxonomic classification and functional annotation, prediction of host range and zoonotic potential, and efforts toward forecasting emerging variants. These advances are transforming viral genomics from a largely descriptive discipline into one with increasing predictive capability. We also critically assess the major challenges that constrain current approaches, including the availability of high-quality and representative datasets, rigorous model evaluation, biological interpretability and responsible governance for increasingly capable AI models.

Artificial Intelligence↗

ORMDL proteins are a conserved new family of endoplasmic reticulum membrane proteins.

BACKGROUND: Annotations of completely sequenced genomes reveal that nearly half of the genes identified are of unknown function, and that some belong to uncharacterized gene families. To help resolve such issues, information can be obtained from the comparative analysis of homologous genes in model organisms. RESULTS: While characterizing genes from the retinitis pigmentosa locus RP26 at 2q31-q33, we have identified a new gene, ORMDL1, that belongs to a novel gene family comprising three genes in humans (ORMDL1, ORMDL2 and ORMDL3), and homologs in yeast, microsporidia, plants, Drosophila, urochordates and vertebrates. The human genes are expressed ubiquitously in adult and fetal tissues. The Drosophila ORMDL homolog is also expressed throughout embryonic and larval stages, particularly in ectodermally derived tissues. The ORMDL genes encode transmembrane proteins anchored in the endoplasmic reticulum (ER). Double knockout of the two Saccharomyces cerevisiae homologs leads to decreased growth rate and greater sensitivity to tunicamycin and dithiothreitol. Yeast mutants can be rescued by human ORMDL homologs. CONCLUSIONS: From protein sequence comparisons we have defined a novel gene family, not previously recognized because of the absence of a characterized functional signature. The sequence conservation of this family from yeast to vertebrates, the maintenance of duplicate copies in different lineages, the ubiquitous pattern of expression in human and Drosophila, the partial functional redundancy of the yeast homologs and phenotypic rescue by the human homologs, strongly support functional conservation. Subcellular localization and the response of yeast mutants to specific agents point to the involvement of ORMDL in protein folding in the ER.

Amino Acid Sequence↗

FANTOM DB: database of Functional Annotation of RIKEN Mouse cDNA Clones.

FANTOM DB, the database of Functional Annotation of RIKEN Mouse cDNA Clones, is designed to store sequence information of RIKEN full-length enriched mouse cDNA clones, graphical views of sequence analysis results, curated functional annotation information and additional descriptions, including Gene Ontology terms. RIKEN's Mouse Gene Encyclopedia Project aims to collect full-length enriched cDNA clones from various mouse tissues, determine the full-length nucleotide sequences, infer their chromosomal locations by computer and characterize gene expression patterns. FANTOM DB has been developed to facilitate this work and to facilitate functional genomic studies such as positional candidate cloning, cDNA microarrays and protein interaction analyses. FANTOM DB contains 21 076 full-length cDNA sequences with rich functional annotations and is publicly available. FANTOM DB thus provides curated functional annotation to RIKEN full-length enriched mouse clones, and has links to other public resources. FANTOM DB can be accessed at http://fantom.gsc.riken.go.jp/db/.

Animals↗

MitoDrome: a database of Drosophila melanogaster nuclear genes encoding proteins targeted to the mitochondrion.

Mitochondria are organelles present in the cytoplasm of most eukaryotic cells; although they have their own DNA, the majority of the proteins necessary for a functional mitochondrion are coded by the nuclear DNA and only after transcription and translation they are imported in the mitochondrion as proteins. The primary role of the mitochondrion is electron transport and oxidative phosphorylation. Although it has been studied for a long time, the interest of researchers in mitochondria is still alive thanks to the discovery of mitochondrial role in apoptosis, aging and cancer. Aim of the MitoDrome database is to annotate the Drosophila melanogaster nuclear genes coding for mitochondrial proteins in order to contribute to the functional characterization of nuclear genes coding for mitochondrial proteins and to knowledge of gene diseases related to mitochondrial dysfunctions. Indeed D. melanogaster is one of the most studied organisms and a model for the Human genome. Data are derived from the comparison of Human mitochondrial proteins versus the Drosophila genome, ESTs and cDNA sequence data available in the FlyBase database. Links from the MitoDrome entries to the related homologous entries available in MitoNuC will be soon imple-mented. The MitoDrome database is available at http://bighost.area.ba.cnr.it/BIG/MitoDrome. Data are organised in a flat-file format and can be retrieved using the SRS system.

Animals↗

InterPro, progress and status in 2005.

InterPro, an integrated documentation resource of protein families, domains and functional sites, was created to integrate the major protein signature databases. Currently, it includes PROSITE, Pfam, PRINTS, ProDom, SMART, TIGRFAMs, PIRSF and SUPERFAMILY. Signatures are manually integrated into InterPro entries that are curated to provide biological and functional information. Annotation is provided in an abstract, Gene Ontology mapping and links to specialized databases. New features of InterPro include extended protein match views, taxonomic range information and protein 3D structure data. One of the new match views is the InterPro Domain Architecture view, which shows the domain composition of protein matches. Two new entry types were introduced to better describe InterPro entries: these are active site and binding site. PIRSF and the structure-based SUPERFAMILY are the latest member databases to join InterPro, and CATH and PANTHER are soon to be integrated. InterPro release 8.0 contains 11 007 entries, representing 2573 domains, 8166 families, 201 repeats, 26 active sites, 21 binding sites and 20 post-translational modification sites. InterPro covers over 78% of all proteins in the Swiss-Prot and TrEMBL components of UniProt. The database is available for text- and sequence-based searches via a webserver (http://www.ebi.ac.uk/interpro), and for download by anonymous FTP (ftp://ftp.ebi.ac.uk/pub/databases/interpro).

Databases, Protein↗