Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,279 records · Page 71Linked to original sources

Distinguishing the ORFs from the ELFs: short bacterial genes and the annotation of genomes.

A substantial fraction of hypothetical open reading frames (ORFs) in completely sequenced bacterial genomes are short, suggesting that many are not genes but random stretches of DNA. Although it is not feasible to authenticate the coding capacity of all such regions experimentally, comparisons of ORFs in related genomes can expose those that encode functional proteins.

Bacteria↗

IgStrand: A universal residue numbering scheme for the immunoglobulin-fold (Ig-fold) to study Ig-proteomes and Ig-interactomes.

The Immunoglobulin fold (Ig-fold) is found in proteins from all domains of life and represents the most populous fold in the human genome, with current estimates ranging from 2 to 3% of protein coding regions. That proportion is much higher in the surfaceome where Ig and Ig-like domains orchestrate cell-cell recognition, adhesion and signaling. The ability of Ig-domains to reliably fold and self-assemble through highly specific interfaces represents a remarkable property of these domains, making them key elements of molecular interaction systems: the immune system, the nervous system, the vascular system and the muscular system. We define a universal residue numbering scheme, common to all domains sharing the Ig-fold in order to study the wide spectrum of Ig-domain variants constituting the Ig-proteome and Ig-Ig interactomes at the heart of these systems. The "IgStrand numbering scheme" enables the identification of Ig structural proteomes and interactomes in and between any species, and comparative structural, functional, and evolutionary analyses. We review how Ig-domains are classified today as topological and structural variants and highlight the "Ig-fold irreducible structural signature" shared by all of them. The IgStrand numbering scheme lays the foundation for the systematic annotation of structural proteomes by detecting and accurately labeling Ig-, Ig-like and Ig-extended domains in proteins, which are poorly annotated in current databases and opens the door to accurate machine learning. Importantly, it sheds light on the robust Ig protein folding algorithm used by nature to form beta sandwich supersecondary structures. The numbering scheme powers an algorithm implemented in the interactive structural analysis software iCn3D to systematically recognize Ig-domains, annotate them and perform detailed analyses comparing any domain sharing the Ig-fold in sequence, topology and structure, regardless of their diverse topologies or origin. The scheme provides a robust fold detection and labeling mechanism that reveals unsuspected structural homologies among protein structures beyond currently identified Ig- and Ig-like domain variants. Indeed, multiple folds classified independently contain a common structural signature, in particular jelly-rolls. Examples of folds that harbor an "Ig-extended" architecture are given. Applications in protein engineering around the Ig-architecture are straightforward based on the universal numbering.

Humans↗

Low-frequency Fourier spectrum for predicting membrane protein types.

Cell membranes are vitally important to living cells. Although the infrastructure of biological membrane is provided by the lipid bilayer, membrane proteins perform most of the specific functions. Knowledge of membrane protein types often provides crucial hints toward determining the function of an uncharacterized membrane protein. With the avalanche of new protein sequences generated in the post-genomic era, it is highly demanded to develop a high throughput tool in identifying the type of newly found membrane proteins according to their primary sequences, so as to timely annotate them for reference usage in both basic research and drug discovery. To realize this, the key is to establish a powerful identifier that can catch their characteristic sequence patterns for different membrane protein types. However, it is not easy because they are buried in a pile of long and complicated sequences. In this paper, based on the concept of the pseudo-amino acid composition [K.C. Chou, PROTEINS: Struct., Funct., Genet. 43 (2001) 246-255], the low-frequency Fourier spectrum analysis is introduced. The merits by doing so are that the sequence pattern information can be more effectively incorporated into a set of discrete components, and that all the existing prediction algorithms can be straightforwardly used on such a formulation for protein samples. High success rates were observed by the re-substitution test, jackknife test, and independent dataset test, indicating that the low-frequency Fourier spectrum approach may become a very useful tool for membrane protein type prediction. The novel approach also holds a high potential for predicting many other attributes of proteins.

Fourier Analysis↗

dcHiChIP: a comprehensive Nextflow-based pipeline for multiscale analysis of chromatin architecture from HiChIP data.

MOTIVATION: Despite the growing use of HiChIP to investigate protein-directed chromatin architecture, a comprehensive and reproducible pipeline for analysing these datasets-from raw reads to multiscale 3D genome features-remains lacking. Existing tools often focus on isolated components, such as loop calling or matrix generation, but fall short in integrating structural annotation, functional enrichment, and spatial modeling within a unified framework. To address this gap, we developed dcHiChIP, a modular, scalable Nextflow-based workflow that streamlines the analysis of HiChIP data, enabling both routine processing and in-depth exploration of chromatin organization and regulatory interactions. RESULTS: dcHiChIP enables robust and reproducible analysis of HiChIP datasets across multiple scales of chromatin architecture. It accepts raw sequencing data as input and generates high-quality loop calls, domain annotations, and 3D genome models. It also performs functional annotation and motif enrichment analyses. Applied to benchmark CTCF HiChIP datasets, dcHiChIP identifies major chromatin architectural features such as TADs/CCDs, A/B compartments, and chromatin stripes, and offers efficient, end-to-end execution with support for batch processing and workflow resumability. AVAILABILITY: dcHiChIP is publicly available on GitHub at https://github.com/SFGLab/dcHiChIP, with documentation at https://sfglab.github.io/dcHiChIP/. The software version used in this study is archived at Zenodo: https://doi.org/10.5281/zenodo.22030542.

Chromatin↗

RNAdb 2.0--an expanded database of mammalian non-coding RNAs.

RNAdb is a comprehensive database of mammalian non-protein-coding RNAs (ncRNAs). There is increasing recognition that ncRNAs play important regulatory roles in multicellular organisms, and there is an expanding rate of discovery of novel ncRNAs as well as an increasing allocation of function. In this update to RNAdb, we provide nucleotide sequences and annotations for tens of thousands of non-housekeeping ncRNAs, including a wide range of mammalian microRNAs, small nucleolar RNAs and larger mRNA-like ncRNAs. Some of these have documented functions and/or expression patterns, but the majority remain of unclear significance, and include PIWI-interacting RNAs, ncRNAs identified from the latest rounds of large-scale cDNA sequencing projects, putative antisense transcripts, as well as ncRNAs predicted on the basis of structural features and alignments. Improvements to the database comprise not only new and updated ncRNA datasets, but also provision of microarray-based expression data and closer interface with more specialized ncRNA resources such as miRBase and snoRNA-LBME-db. To access RNAdb, visit http://research.imb.uq.edu.au/RNAdb.

Animals↗

SAGE of the developing wheat caryopsis.

Understanding the development of the cereal caryopsis holds the future for metabolic engineering in the interests of enhancing global food production. We have developed a Serial Analysis of Gene Expression (SAGE) data platform to investigate the developing wheat (Triticum aestivum) caryopsis. LongSAGE libraries have been constructed at five time-points post-anthesis to coincide with key processes in caryopsis development. More than 90,000 LongSAGE tags have been sequenced generating 29,261 unique tag sequences across all five libraries. Tag abundance, generated from cumulative tag counts, provides insight into the redundancy and diversity of each library. Annotation of the 500 most abundant tags spanning development highlights the array of functional groups being expressed. The relative frequency of these more abundant transcripts allows quantitative analysis of patterns of expression during grain development. We have identified activities of cellular proliferation/differentiation, the accumulation of storage proteins and starch biosynthesis. The abundance of calcium-dependent protein kinases indicate their importance in signalling across development. Acquisition of a broad array of defence coincides with storage accumulation and is dominated by inhibitors of amylase activity. Differential expression profiles of abundant tags from each library reveal the coordinated expression of genes responsible for the cellular events constituting caryopsis development. This SAGE platform has also provided a resource of novel sequence and expression information including the identification of potentially useful promoter activities. Further investigations into both the abundant and low expressing transcripts will provide greater insight into wheat caryopsis development and assist in wheat improvement programmes.

Gene Expression Profiling↗

Construction, analysis, and beta-glucanase screening of a bacterial artificial chromosome library from the large-bowel microbiota of mice.

A metagenomic (community genomic) library consisting of 5,760 bacterial artificial chromosome clones was prepared in Escherichia coli DH10B from DNA extracted from the large-bowel microbiota of BALB/c mice. DNA inserts detected in 61 randomly chosen clones averaged 55 kbp (range, 8 to 150 kbp) in size. A functional screen of the library for beta-glucanase activity was conducted using lichenin agar plates and Congo red solution. Three clones with beta-glucanase activity were detected. The inserts of these three clones were sequenced and annotated. Open reading frames (ORF) that encoded putative proteins with identity to glucanolytic enzymes (lichenases and laminarinases) were detected by reference to databases. Other putative genes were detected, some of which might have a role in environmental sensing, nutrient acquisition, or coaggregation. The insert DNA from two clones probably originated from uncultivated bacteria because the ORF had low sequence identity with database entries, but the genes associated with the remaining clone resembled sequences reported in Bacteroides species.

Amino Acid Sequence↗

ORFer--retrieval of protein sequences and open reading frames from GenBank and storage into relational databases or text files.

BACKGROUND: Functional genomics involves the parallel experimentation with large sets of proteins. This requires management of large sets of open reading frames as a prerequisite of the cloning and recombinant expression of these proteins. RESULTS: A Java program was developed for retrieval of protein and nucleic acid sequences and annotations from NCBI GenBank, using the XML sequence format. Annotations retrieved by ORFer include sequence name, organism and also the completeness of the sequence. The program has a graphical user interface, although it can be used in a non-interactive mode. For protein sequences, the program also extracts the open reading frame sequence, if available, and checks its correct translation. ORFer accepts user input in the form of single or lists of GenBank GI identifiers or accession numbers. It can be used to extract complete sets of open reading frames and protein sequences from any kind of GenBank sequence entry, including complete genomes or chromosomes. Sequences are either stored with their features in a relational database or can be exported as text files in Fasta or tabulator delimited format. The ORFer program is freely available at http://www.proteinstrukturfabrik.de/orfer. CONCLUSION: The ORFer program allows for fast retrieval of DNA sequences, protein sequences and their open reading frames and sequence annotations from GenBank. Furthermore, storage of sequences and features in a relational database is supported. Such a database can supplement a laboratory information system (LIMS) with appropriate sequence information.

Animals↗

Molecular characterization of the developmental gene in eyes: through data-mining on integrated transcriptome databases.

OBJECTIVES: Our aim was to utilize publicly available and proprietary sources to discover candidate genes important for ocular development. DESIGN AND METHODS: The collated information on our 5092 non-redundant clusters was grouped and functional annotation was conducted using gene ontology (FatiGO) for categorizing them with respect to molecular function. The web-based viewer technological platform (H-InvDB) was employed for transcription analyses of in-house high quality fetal eye Expressed Sequence Tags (ESTs). Eye-specific ESTs were also analyzed across species by using EMBEST. RESULTS: According to adult eye cDNA libraries, nucleic acid binding and cell structure/cytoskeletal protein genes were the most abundant among the ESTs of fetal eyes. Using cDNA assembly in H-InvDB, 20 (80%) of the 25 most commonly expressed genes in the human eye are also expressed in extraocular tissues. The crystalline gamma S gene is highly expressed in the eye, but not in other tissues. We used EMBEST to compare human fetal eye and octopus eye ESTs and the expression similarity was low (1.6%). This indicated that our fetal eye library contains genes necessary for the developmental process and biological function of the eye, which may not be expressed in the fully developed octopus eyes. The human fetal eye cDNA library also contained highly abundant eye tissue genes, including alphaA-crystallin, eukaryotic translation elongation factor 1 alpha 1 (EEF1A1), bestrophin (VMD2), cystatin C, and transforming growth factor, beta-induced (BIGH3). CONCLUSIONS: Our annotated EST set provides a valuable resource for gene discovery and functional genomic analysis. This display will help to appreciate the strengths and weaknesses of the different technological platforms, so that in future studies the maximum amount of beneficial information can be derived from the appropriate use of each method.

Animals↗

The KEGG databases at GenomeNet.

The Kyoto Encyclopedia of Genes and Genomes (KEGG) is the primary database resource of the Japanese GenomeNet service (http://www.genome.ad.jp/) for understanding higher order functional meanings and utilities of the cell or the organism from its genome information. KEGG consists of the PATHWAY database for the computerized knowledge on molecular interaction networks such as pathways and complexes, the GENES database for the information about genes and proteins generated by genome sequencing projects, and the LIGAND database for the information about chemical compounds and chemical reactions that are relevant to cellular processes. In addition to these three main databases, limited amounts of experimental data for microarray gene expression profiles and yeast two-hybrid systems are stored in the EXPRESSION and BRITE databases, respectively. Furthermore, a new database, named SSDB, is available for exploring the universe of all protein coding genes in the complete genomes and for identifying functional links and ortholog groups. The data objects in the KEGG databases are all represented as graphs and various computational methods are developed to detect graph features that can be related to biological functions. For example, the correlated clusters are graph similarities which can be used to predict a set of genes coding for a pathway or a complex, as summarized in the ortholog group tables, and the cliques in the SSDB graph are used to annotate genes. The KEGG databases are updated daily and made freely available (http://www.genome.ad.jp/kegg/).

Animals↗

In planta horizontal transfer of a major pathogenicity effector gene.

Xanthomonas citri pv. citri is a clonal group of strains that causes citrus canker disease and appears to have originated in Asia. A phylogenetically distinct clonal group that causes identical disease symptoms on susceptible citrus, X. citri pv. aurantifolii, arose more recently in South America. Genomes of X. citri pv. aurantifolii strains carry two DNA fragments that hybridize to pthA, an X. citri pv. citri gene which encodes a major type III pathogenicity effector protein that is absolutely required to cause citrus canker. Marker interruption mutagenesis and complementation revealed that X. citri pv. aurantifolii strain B69 carried one functional pthA homolog, designated pthB, that was required to cause cankers on citrus. Gene pthB was found among 38 open reading frames on a 37,106-bp plasmid, designated pXcB, which was sequenced and annotated. No additional pathogenicity effectors were found on pXcB, but 11 out of 38 open reading frames appeared to encode a type IV transfer system. pXcB transferred horizontally in planta, without added selection, from B69 to a nonpathogenic X. citri pv. citri (pthA::Tn5) mutant strain, fully restoring canker. In planta transfer efficiencies were very high (>0.1%/recipient) and equivalent to those observed for agar medium with antibiotic selection, indicating that pthB conferred a strong selective advantage to the recipient strain. A single pathogenicity effector that can confer a distinct selective advantage in planta may both facilitate plasmid survival following horizontal gene transfer and account for the origination of phylogenetically distinct groups of strains causing identical disease symptoms.

Bacterial Proteins↗

Predictive screening for regulators of conserved functional gene modules (gene batteries) in mammals.

BACKGROUND: The expression of gene batteries, genomic units of functionally linked genes which are activated by similar sets of cis- and trans-acting regulators, has been proposed as a major determinant of cell specialization in metazoans. We developed a predictive procedure to screen the mouse and human genomes and transcriptomes for cases of gene-battery-like regulation. RESULTS: In a screen that covered approximately 40 percent of all annotated protein-coding genes, we identified 21 co-expressed gene clusters with statistically supported sharing of cis-regulatory sequence elements. 66 predicted cases of over-represented transcription factor binding motifs were validated against the literature and fell into three categories: (i) previously described cases of gene battery-like regulation, (ii) previously unreported cases of gene battery-like regulation with some support in a limited number of genes, and (iii) predicted cases that currently lack experimental support. The novel predictions include for example Sox 17 and RFX transcription factor binding sites that were detected in approximately 10% of all testis specific genes, and HNF-1 and 4 binding sites that were detected in approximately 30% of all kidney specific genes respectively. The results are publicly available at http://www.wlab.gu.se/lindahl/genebatteries. CONCLUSION: 21 co-expressed gene clusters were enriched for a total of 66 shared cis-regulatory sequence elements. A majority of these predictions represent novel cases of potential co-regulation of functionally coupled proteins. Critical technical parameters were evaluated, and the results and the methods provide a valuable resource for future experimental design.

Amino Acid Motifs↗

Stability of plasmid pA387 derivatives in Amycolatopsis mediterranei producing rifamycin.

Genetic studies on the biosynthesis of rifamycins in producer strains such as Amylcolaptopsis mediterranei U-32 are severely hampered by the availability of efficient transformation procedures and stable plasmid vectors. Using an efficient electroporation procedure we have studied the replication and stability of a pA387 derivative, pDXM32. This plasmid confers enhanced plasmid stability and copy number compared to pA387 derivatives commonly used as cloning vectors in A. mediterranei. Deletion derivatives in the region previously identified as being a minimal replication origin were also examined with respect to their ability to transform A. mediterranei and at least one locus was essential for replication. A 5.4 kbp DNA fragment was sequenced and annotated encoding the replication and plasmid stability functions. A parA homologue was identified which is likely to confer plasmid stability.

Actinomycetaceae↗

An organism-specific method to rank predicted coding regions in Trypanosoma brucei.

Genome annotation in differently evolved organisms presents challenges because the lack of sequence-based homology limits the ability to determine the function of putative coding regions. To provide an alternative to annotation by sequence homology, we developed a method that takes advantage of unusual trypanosomatid biology and skews in nucleotide composition between coding regions and upstream regions to rank putative open reading frames based on the likelihood of coding. The method is 93% accurate when tested on known genes. We have applied our method to the full complement of open reading frames on Chromosome I of Trypanosoma brucei, and we can predict with high confidence that 226 putative coding regions are likely to be functional. Methods such as the one described here for discriminating true coding regions are critical for genome annotation when other sources of evidence for function are limited.

Animals↗

Two large Arabidopsis thaliana gene families are homologous to the Brassica gene superfamily that encodes pollen coat proteins and the male component of the self-incompatibility response.

The male component of the self-incompatibility response in Brassica has recently been shown to be encoded by the S locus cysteine-rich gene (SCR). SCR is related, at the sequence level, to the pollen coat protein (PCP) gene family whose members encode small, cysteine-rich proteins located in the proteo-lipidic surface layer (tryphine) of Brassica pollen grains. Here we show that the Arabidopsis genome includes two large gene families with homology to SCR and to the PCP gene family, respectively. These genes are poorly predicted by gene-identification algorithms and, with few exceptions, have been missed in previous annotations. Based on sequence comparison and an analysis of the expression patterns of several members of each family, we discuss the possible functions of these genes. In particular, we consider the possibility that SCR-related genes in Arabidopsis may encode ligands for the S gene family of receptor-like kinases in this species.

Alleles↗

Analogous enzymes: independent inventions in enzyme evolution.

It is known that the same reaction may be catalyzed by structurally unrelated enzymes. We performed a systematic search for such analogous (as opposed to homologous) enzymes by evaluating sequence conservation among enzymes with the same enzyme classification (EC) number using sensitive, iterative sequence database search methods. Enzymes without detectable sequence similarity to each other were found for 105 EC numbers (a total of 243 distinct proteins). In 34 cases, independent evolutionary origin of the suspected analogous enzymes was corroborated by showing that they possess different structural folds. Analogous enzymes were found in each class of enzymes, but their overall distribution on the map of biochemical pathways is patchy, suggesting multiple events of gene transfer and selective loss in evolution, rather than acquisition of entire pathways catalyzed by a set of unrelated enzymes. Recruitment of enzymes that catalyze a similar but distinct reaction seems to be a major scenario for the evolution of analogous enzymes, which should be taken into account for functional annotation of genomes. For many analogous enzymes, the bacterial form of the enzyme is different from the eukaryotic one; such enzymes may be promising targets for the development of new antibacterial drugs.

Amino Acid Sequence↗

pdb-care (PDB carbohydrate residue check): a program to support annotation of complex carbohydrate structures in PDB files.

BACKGROUND: Carbohydrates are involved in a variety of fundamental biological processes and pathological situations. They therefore have a large pharmaceutical and diagnostic potential. Knowledge of the 3D structure of glycans is a prerequisite for a complete understanding of their biological functions. The largest source of biomolecular 3D structures is the Protein Data Bank. However, about 30% of all 1663 PDB entries (version September 2003) containing carbohydrates comprise errors in glycan description. Unfortunately, no software is currently available which aligns the 3D information with the reported assignments. It is the aim of this work to fill this gap. RESULTS: The pdb-care program http://www.glycosciences.de/tools/pdb-care/ is able to identify and assign carbohydrate structures using only atom types and their 3D atom coordinates given in PDB-files. Looking up a translation table where systematic names and the respective PDB residue codes are listed, both assignments are compared and inconsistencies are reported. Additionally, the reliability of reported and calculated connectivities for molecules listed within the HETATOM records is checked and unusual values are reported. CONCLUSION: Frequent use of pdb-care will help to improve the quality of carbohydrate data contained in the PDB. Automatic assignment of carbohydrate structures contained in PDB entries will enable the cross-linking of glycobiology resources with genomic and proteomic data collections.

Carbohydrate Conformation↗

Genome assembly and annotation of the parasitoid jewel wasp Nasonia oneida.

The jewel wasp, Nasonia (Hymenoptera: Pteromalidae), is a well-established model system for evolutionary genetics and host-microbial interactions. Here, we present the genome of N. oneida, a species lacking prior genomic characterization, using 10× Genomics linked-read (400× coverage), Illumina short-read (120× coverage), and transcriptome data (30× coverage). The assembled genome size is 267 Mb, comprising 4,675 scaffolds, with a scaffold N50 of 1 Mb and 98.40% Benchmarking Universal Single-Copy Orthologues (BUSCOs) completeness score. Annotation revealed 32.29% (86.46 Mb) of repetitive sequences and 14,221 protein-coding genes. Comparative genomics of N. oneida with 15 other hymenopteran species validated the presence of 5,939 gene families shared among them, including 3643 single-copy and 2296 multicopy gene families. This study provides the first de novo assembly of N. oneida, providing a significant addition to the growing repertoire of molecular tools for comparative genomics and functional studies to understand the evolution of closely related species as well as the evolution of parasitic wasps.

Animals↗