Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 955 records · Page 53Linked to original sources

Tox-Prot, the toxin protein annotation program of the Swiss-Prot protein knowledgebase.

The Tox-Prot program was initiated in order to provide the scientific community a summary of the current knowledge on animal protein toxins. The aim of this program is to systematically annotate all proteins which act as toxins and are produced by venomous and poisonous animals. Venomous animals such as snakes, scorpions, spiders, jellyfish, insects, cone snails, sea anemones, lizards, some fish, and platypus are equipped with a specialized organ to inject venom in their prey. In contrast, poisonous animals such as some fish or worms, lack such organs. Each toxin is annotated according to the quality standards of Swiss-Prot. This means providing a wealth of information that includes the description of the function, domain structure, subcellular location, tissue specificity, variants, similarities to other proteins, keywords, etc. In the framework of this program, particular care has been made to capture what is known on the function and mode of action, posttranslational modifications and 3D structural data which are all relatively abundant in the field of protein toxins. Researchers are welcome to contribute their knowledge to the scientific community by submitting relevant findings to Swiss-Prot concerning toxins at Tox-Prot@isb-sib.ch. More information on Tox-Prot can be found at http://www.expasy.org/sprot/tox-prot.

Animals↗

Identification of function-associated loop motifs and application to protein function prediction.

MOTIVATION: The detection of function-related local 3D-motifs in protein structures can provide insights towards protein function in absence of sequence or fold similarity. Protein loops are known to play important roles in protein function and several loop classifications have been described, but the automated identification of putative functional 3D-motifs in such classifications has not yet been addressed. This identification can be used on sequence annotations. RESULTS: We evaluated three different scoring methods for their ability to identify known motifs from the PROSITE database in ArchDB. More than 500 new putative function-related motifs not reported in PROSITE were identified. Sequence patterns derived from these motifs were especially useful at predicting precise annotations. The number of reliable sequence annotations could be increased up to 100% with respect to standard BLAST. CONTACT: boliva@imim.es SUPPLEMENTARY INFORMATION: Supplementary Data are available at Bioinformatics online.

Amino Acid Sequence↗

Annotating the human proteome.

The completion of the human genome has shifted the attention from deciphering the sequence to the identification and characterization of the encoded components. The identification and functional annotation of the proteome is here of special interest and starts with the identification of genes and transcripts as a prerequisite of proteome annotation. Gene predictions are very powerful in predicting most of the exons in a genome, but reliable gene structure predictions of both known and novel genes are dependent on existing transcript and protein information. An enormous amount of data already exists on the function of many human proteins, but this is scattered over many resources. Public domain databases are required to manage and collate this information and present it to the user community in both a human and machine readable manner.

Databases, Factual↗

Genome analysis of the glycosphingolipid-producing green alga tetraselmis sp. NKG400013.

Microalgae are gaining attention as sustainable resources for the production of valuable compounds, including biofuels, pigments, and bioactive metabolites. To support metabolic engineering and genome editing approaches aimed at enhancing these traits, high-quality genome assemblies are essential; however, genomic information remains limited for many microalgal lineages. Tetraselmis sp. NKG400013 is a green alga known for high glycosphingolipid accumulation with distinctive structural features. Here, we report a draft genome assembly of this strain generated using PacBio HiFi sequencing and transcriptome-supported annotation. The assembled genome spans 423.7 Mbp, with 74.5% repetitive sequences and 15,322 predicted protein-coding genes. Comparative analyses across 11 green algal species revealed a positive correlation between genome sizes and repeat contents, indicating that transposable element expansion, particularly long terminal repeat retrotransposons, has substantially contributed to genome enlargement in Tetraselmis. Genome-wide functional annotation and ortholog inference identified core enzymes required for glycosylceramide biosynthesis. Both sphingolipid Δ4 and Δ8 desaturases were identified in Tetraselmis and their coexistence suggests an expanded capacity for long-chain base modification that may underlie its distinctive glycosphingolipid profile. These results establish a genomic framework for understanding the high glycosphingolipid-producing capacity of NKG400013 and provide insights into the evolutionary diversification of sphingolipid metabolism in green algae.

Chlorophyta↗

DirectASRM: uncovering allele-specific post-transcriptional RNA modifications through direct RNA sequencing.

SUMMARY: We developed DirectASRM, a comprehensive database for the systematic identification, integration, and annotation of allele-specific RNA modifications (ASRMs) from direct RNA sequencing data. DirectASRM enables single-base, transcript-level detection of ASRMs across multiple RNA modification types, diverse organisms and condition-specific contexts. The database further evaluates the confidence of each ASRM-SNP pair association within isoform context by jointly considering statistical evidence of allelic modification imbalance and independent support from external next-generation sequencing (NGS) - based RNA modification resources. DirectASRM also provides extensive functional annotations for ASRMs and their associated variants, including intra-sample transcript-level allele-specific expression (ASE) and allele-specific splicing, as well as additional post-transcriptional regulatory features such as miRNA binding, circRNA, RNA-protein interactions, and disease relevance. Overall, DirectASRM serves as a comprehensive resource that supports systematic investigation of the potential functional impact of genetic variants in epitranscriptomic regulation. AVAILABILITY AND IMPLEMENTATION: DirectASRM database is freely accessible at http://modinfor.com/DirectASRM/. DirectASRM pipeline is available at GitHub (https://github.com/jiayin1101/DirectASRM_pipeline) and Zenodo (DOI: https://doi.org/10.5281/zenodo.19876077).

Alleles↗

PANDORA: keyword-based analysis of protein sets by integration of annotation sources.

Recent advances in high-throughput methods and the application of computational tools for automatic classification of proteins have made it possible to carry out large-scale proteomic analyses. Biological analysis and interpretation of sets of proteins is a time-consuming undertaking carried out manually by experts. We have developed PANDORA (Protein ANnotation Diagram ORiented Analysis), a web-based tool that provides an automatic representation of the biological knowledge associated with any set of proteins. PANDORA uses a unique approach of keyword-based graphical analysis that focuses on detecting subsets of proteins that share unique biological properties and the intersections of such sets. PANDORA currently supports SwissProt keywords, NCBI Taxonomy, InterPro entries and the hierarchical classification terms from ENZYME, SCOP and GO databases. The integrated study of several annotation sources simultaneously allows a representation of biological relations of structure, function, cellular location, taxonomy, domains and motifs. PANDORA is also integrated into the ProtoNet system, thus allowing testing thousands of automatically generated clusters. We illustrate how PANDORA enhances the biological understanding of large, non-uniform sets of proteins originating from experimental and computational sources, without the need for prior biological knowledge on individual proteins.

Computational Biology↗

FlgM anti-sigma factors: identification of novel members of the family, evolutionary analysis, homology modeling, and analysis of sequence-structure-function relationships.

FlgM proteins, also known as Anti-sigma-28 factor (sigma28), are negative regulators of flagellin synthesis. Recently, a three-dimensional structure of the Aquifex aeolicus sigma28/FlgM complex (PDB code: 1rp3) was determined by X-ray crystallography at 2.3 A resolution. Furthermore, experimental data on bacterial FlgM, including site-directed mutagenesis and structural characterization by NMR are also available. However, an interpretation of the sequence-structure-function relationships combining X-ray and NMR data with the evolutionary information extracted from the increasing number of FlgM-related sequences annotated in databases is not available. In the present study, we combined database sequence searches and sequence-analysis tools to update the multiple sequence alignment of a previously characterized cluster of orthologs (COG2747) and the PFAM classification of protein domains (PF04316) for the FlgM family. A phylogenetic analysis of 77 protein sequences revealed the presence of at least three major sequence clades within the FlgM family. Besides, we predicted functional residues using a SequenceSpace method. We also generated homology models for Bacillus subtilis and Salmonella typhimurium FlgM proteins, for which sequence-structure-function relationship data are available, and used the docking program ClusPro to hypothesize about the dimer association between FlgM proteins. In conclusion, the analysis presented in this work will be useful in designing new experiments to understand better protein-protein interactions between FglM, sigma factors, and putative molecules from the flagellar export apparatus. Electronic Supplementary Material is available in the online version of this article at http://link.springer.de/

Bacterial Proteins↗

Open reading frames provide a rich pool of potential natural antisense transcripts in fungal genomes.

Natural antisense transcripts are reported from all kingdoms of life and several recent reports of genomewide screens indicate that they are widely distributed. These transcripts seem to be involved in various biological functions and may govern the expression of their respective sense partner. Very little, however, is known about the degree of evolutionary conservation of antisense transcripts. Furthermore, none of the earlier analyses has studied whether antisense relationships are solely dual or involved in more complex relationships. Here we present a systematic screen for cis- and trans-located antisense transcripts based on open reading frames (ORFs) from five fungal species. The relative number of ORFs involved in antisense relationships varies greatly between the five species. In addition, other significant differences are found between the species, such as the mean length of the antisense region. The majority of trans-located antisense transcripts is found to be involved in complex relationships, resulting in highly connected networks. The analysis of the degree of evolutionary conservation of antisense transcripts shows that most antisense transcripts have no ortholog in any other species. An annotation of antisense transcripts based on Gene Ontology directs to common terms and shows that proteins of genes involved in antisense relationships preferentially localize to the nucleus with common functions in the regulation or maintenance of nucleic acids.

Evolution, Molecular↗

Dynamic transcriptome of mice.

Life science in the 21st century is developing rapidly through the structural analysis of biomolecules, the completion of the human genome sequence and the analysis of transcriptomes. The mouse transcriptome has been comprehensively analyzed using a gene discovery approach to collect full-length cDNA (FL-cDNA) clones. The framework of the transcriptome was then mapped out by an international Functional ANnoTation Of Mouse cDNA (FANTOM) effort, and a significant new population of noncoding transcripts was discovered. The geographical analogy of a second "RNA continent," separate from the "continent" of expressed proteins, aids the visualization of this concept. An unexpected number of variations was discovered in the mouse transcriptome. The animal transcriptome has evolved to produce several transcripts and proteins from a single "transcriptional unit". Transcriptome analysis has given rise to the FL-cDNA database and to the 60 770 FANTOM FL-cDNA clone set, and the DNABook was developed as an easier way to distribute these clones. In conjunction with genome sequence databases, transcriptome databases and clone banks will be platforms for developing advanced databases of gene function (e.g. the Genome Function Database). This will enable life science to make rapid progress towards understanding life as a system of molecules.

Animals↗

Identification of novel clock-controlled genes by cDNA macroarray analysis in Chlamydomonas reinhardtii.

Circadian rhythms are self-sustaining oscillations whose period length under constant conditions is about 24 h. Circadian rhythms are widespread and involve functions as diverse as human sleep-wake cycles and cyanobacterial nitrogen fixation. In spite of a long research history, knowledge about clock-controlled genes is limited in Chlamydomonas reinhardtii. Using a cDNA macroarray containing 10 368 nuclear-encoded genes, we examined global circadian regulation of transcription in Chlamydomonas. We identified 269 candidates for circadianly expressed gene. Northern blot analysis confirmed reproducible and sustainable rhythmicity for 12 genes. Most genes exhibited peak expression at the transition point between day and night. One hundred and eighteen genes were assigned predicted annotations. The functions of the cycling genes were diverse and included photosynthesis, respiration, cellular structure, and various metabolic pathways. Surprisingly, 18 genes encoding chloroplast ribosomal proteins showed a coordinated circadian pattern of expression and peaked just at the beginning of subjective day. The co-regulation of genes bearing a similar function was also observed in genes involved in cellular structure. They peaked at the end of the subjective night, which is when the regeneration of cell walls and flagella in daughter cells occurs. Expression of the chlamyopsin gene, which encodes an opsin-type photoreceptor, also exhibited circadian rhythm.

Animals↗

Probabilistic alignment of motifs with sequences.

MOTIVATION: Motif detection is an important component of the classification and annotation of protein sequences. A method for aligning motifs with an amino acid sequence is introduced. The motifs can be described by the secondary (i.e. functional, biophysical, etc.) characteristics of a signal or pattern to be detected. The results produced are based on the statistical relevance of the alignment. The method was targeted to avoid the problems (i.e. over-fitting, biological interpretation and mathematical soundness) encountered in other methods currently available. RESULTS: The method was tested on lipoprotein signals in B. subtilis yielding stable results. The results of signal prediction were consistent with other methods where literature was available. AVAILABILITY: An implementation of the motif alignment, refining and bootstrapping is available for public use online at http://www.expasy.org/tools/patoseq/

Amino Acid Motifs↗

Single nucleotide polymorphisms associated with rat expressed sequences.

Single nucleotide polymorphisms (SNPs) are the most common source of genetic variation in populations and are thus most likely to account for the majority of phenotypic and behavioral differences between individuals or strains. Although the rat is extensively studied for the latter, data on naturally occurring polymorphisms are mostly lacking. We have used publicly available sequences consisting of whole-genome shotgun (WGS), expressed sequence tag (EST), and mRNA data as a source for the in silico identification of SNPs in gene-coding regions and have identified a large collection of 33,305 high-quality candidate SNPs. Experimental verification of 471 candidate SNPs using a limited set of rat isolates revealed a confirmation rate of approximately 50%. Although the majority of SNPs were identified between Sprague-Dawley (EST data) and Brown Norway (WGS data) strains, we found that 66% of the verified variations are common among different rat strains. All SNPs were extensively annotated, including chromosomal and genetic map information, and nonsynonymous SNPs were analyzed by SIFT and PolyPhen prediction programs for their potential deleterious effect on protein function. Interestingly, we retrieved three SNPs from the database that result in the introduction of a premature stop codon and that could be confirmed experimentally. Two of these "in silico-identified knockouts" reside in interesting QTL regions. Data are publicly available via a Web interface (http://cascad.niob.knaw.nl), allowing simple and advanced search queries.

Animals↗

Recognition of transmembrane segments in proteins: review and consistency-based benchmarking of internet servers.

Membrane proteins perform a number of crucial functions as transporters, receptors, and components of enzyme complexes. Identification of membrane proteins and prediction of their topology is thus an important part of genome annotation. We present here an overview of transmembrane segments in protein sequences, summarize data from large-scale genome studies, and report results of benchmarking of several popular internet servers.

Algorithms↗

Multidimensional protein identification technology: current status and future prospects.

Protein profiling using high-throughput tandem mass spectrometry has become a powerful method for analyzing changes in global protein expression patterns in cells and tissues as a function of developmental, physiologic and disease processes. This review summarizes the utility and practical application of multidimensional protein identification technology as a platform for comprehensive proteomic profiling of complex biologic samples. The strengths and potential problems and limitations associated with this powerful technology are discussed, with an emphasis placed on one of the biggest challenges currently facing large-scale expression profiling projects -- namely, data analysis. Complementary bioinformatic computational data mining strategies, such as clustering, functional annotation and statistical inference, are also discussed as these are increasingly necessary for interpreting the results of global proteomic profiling studies.

Animals↗

Comprehensive identification of Drosophila dorsal-ventral patterning genes using a whole-genome tiling array.

Dorsal-ventral (DV) patterning of the Drosophila embryo is initiated by Dorsal, a sequence-specific transcription factor distributed in a broad nuclear gradient in the precellular embryo. Previous studies have identified as many as 70 protein-coding genes and one microRNA (miRNA) gene that are directly or indirectly regulated by this gradient. A gene regulation network, or circuit diagram, including the functional interconnections among 40 Dorsal target genes and 20 associated tissue-specific enhancers, has been determined for the initial stages of gastrulation. Here, we attempt to extend this analysis by identifying additional DV patterning genes using a recently developed whole-genome tiling array. This analysis led to the identification of another 30 protein-coding genes, including the Drosophila homolog of Idax, an inhibitor of Wnt signaling. In addition, remote 5' exons were identified for at least 10 of the approximately 100 protein-coding genes that were missed in earlier annotations. As many as nine intergenic uncharacterized transcription units were identified, including two that contain known microRNAs, miR-1 and -9a. We discuss the potential functions of these recently identified genes and suggest that intronic enhancers are a common feature of the DV gene network.

Animals↗

Identification and distribution of protein families in 120 completed genomes using Gene3D.

Using a new protocol, PFscape, we undertake a systematic identification of protein families and domain architectures in 120 complete genomes. PFscape clusters sequences into protein families using a Markov clustering algorithm (Enright et al., Nucleic Acids Res 2002;30:1575-1584) followed by complete linkage clustering according to sequence identity. Within each protein family, domains are recognized using a library of hidden Markov models comprising CATH structural and Pfam functional domains. Domain architectures are then determined using DomainFinder (Pearl et al., Protein Sci 2002;11:233-244) and the protein family and domain architecture data are amalgamated in the Gene3D database (Buchan et al., Genome Res 2002;12:503-514). Using Gene3D, we have investigated protein sequence space, the extent of structural annotation, and the distribution of different domain architectures in completed genomes from all kingdoms of life. As with earlier studies by other researchers, the distribution of domain families shows power-law behavior such that the largest 2,000 domain families can be mapped to approximately 70% of nonsingleton genome sequences; the remaining sequences are assigned to much smaller families. While approximately 50% of domain annotations within a genome are assigned to 219 universal domain families, a much smaller proportion (< 10%) of protein sequences are assigned to universal protein families. This supports the mosaic theory of evolution whereby domain duplication followed by domain shuffling gives rise to novel domain architectures that can expand the protein functional repertoire of an organism. Functional data (e.g. COG/KEGG/GO) integrated within Gene3D result in a comprehensive resource that is currently being used in structure genomics initiatives and can be accessed via http://www.biochem.ucl.ac.uk/bsm/cath/Gene3D/.

Amino Acid Sequence↗

Characterization of the type III secretion locus of Bordetella pertussis.

Multiple sequence comparisons of proteins of the LcrD/FlbF family allowed the design of primers that specifically amplify sequences coding for type III secretion components. Amplification of Bordetella pertussis DNA with these primers yielded a fragment that was further used as a probe for screening a genomic library. The nucleotide sequence of a positive clone revealed a 2100-bp gene, called bcrD, which specifies a 75-kDa polypeptide homologous to the Yersinia LcrD protein. Chromosome walking allowed the characterization of a 35-kb DNA segment that contains the entire locus and flanking housekeeping genes. The B. pertussis type III secretion locus consists of more than 30 open reading frames (ORFs), most of which are identical to annotated genes of Bordetella spp and share similarities with known type III secretion genes of related bacteria. In order to assess the function of this locus, we engineered a bcrD null mutant. However, none of the tested phenotypes, such as protein secretion, cellular invasion, cytotoxicity or mouse lung colonization, differentiated the mutant from its parental strain. Studies of bcrD and bscN expressions indicated that, under our experimental conditions, these genes are not expressed in vitro. Restriction analyses on pulsed-field gel electrophoresis allowed the type III locus mapping at coordinate position 1,590 kb on the Tohama I strain chromosome.

Animals↗

Functional characterization of ketoreductase (rubN6) and aminotransferase (rubN4) genes in the gene cluster of Streptomyces achromogenes var. rubradiris.

ORF's for rubN6 and rubN4 have been annotated as thymidine diphosphate glucose 4-ketoreductase and thymidine diphosphate glucose 3-aminotransferase by sequence analysis of the rubradirin biosynthetic gene cluster cloned from Streptomyces achromogenes var. rubradiris NRRL 3061. Both ORFs were heterologously expressed in Escherichia coli as His-tagged fusion proteins. The functionalities of TDP-glucose 4-ketoreductase and TDP-glucose 3-aminotransferase were verified by in vitro enzyme assay, and a biosynthetic pathway for TDP-D: -rubranitrose is proposed.

Amino Acid Sequence↗