Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,747 records · Page 97Linked to original sources

Expression profile of the channel catfish spleen: analysis of genes involved in immune functions.

Both qualitative and quantitative patterns of tissue-specific gene expression can be determined using gene profiling. Expressed sequence tag (EST) analysis is an efficient approach not only for gene discovery and examining gene expression, but also for development of molecular resources useful for functional genomics. As part of an ongoing transcriptome analysis of channel catfish (Ictalurus punctatus), EST analysis was conducted for gene annotations and profiling using a complementary DNA library developed from messenger RNA of the spleen. A total of 1204 spleen cDNA clones were analyzed. Of the 1204 clones, 665 clones (55.2%) were identified as orthologs of known genes from other organisms by BLAST searches and 539 clones (44.8%) as unknown gene clones. In total 147 novel genes were identified, and annotations were made to 118 of them. In addition, 389 novel EST clusters were identified. Expression profile was analyzed in relation to metabolic functional groups. A total of 28 known genes were involved in immune functions, of which 10 were identified for the first time in channel catfish. Microsatellite-containing clones were also identified that may be potentially useful for genome mapping. This work contributed to the Catfish Gene Index, and toward a Unigene set useful for functional genomics research concerning spleen gene functions in relation to disease defenses.

Journal Article↗

Nucleosome formation potential of eukaryotic DNA: calculation and promoters analysis.

MOTIVATION: A rapid growth in the number of genes with known sequences calls for developing automated tools for their classification and analysis. It became clear that nucleosome packaging of eukaryotic DNA is very important for gene functioning. Automated computer tools for characterization of nucleosome packaging density could be useful for studying of gene regulation and genome annotation. RESULTS: A program for constructing nucleosome formation potential profiles of eukaryotic DNA sequences was developed. Nucleosome packaging density was analyzed for different functional types of human promoters. It was found that in promoters of tissue-specific genes, the nucleosome formation potential was essentially higher than in genes expressed in many tissues, or housekeeping genes. Hence, capability of nucleosome positioning in the promoter region may serve as a factor regulating gene expression. AVAILABILITY: The program for nucleosome sites recognition is included into the GeneExpress system; section 'DNA Nucleosomal Organization', http://wwwmgs.bionet.nsc.ru/mgs/programs/recon/.

Algorithms↗

Phylogenetic comparison of metabolic capacities of organisms at genome level.

Horizontal gene transfer (HGT) has been shown to widely spread in organisms by comparative genomic studies. However, its effect on the phylogenetic relationship of organisms, especially at a system level of different cellular functions, is still not well understood. In this work, we have constructed phylogenetic trees based on the enzyme, reaction, and gene contents of metabolic networks reconstructed from annotated genome information of 82 sequenced organisms. Results from different phylogenetic distance definitions and based on three different functional subsystems (i.e., metabolism, cellular processes, information storage and processing) were compared. Results based on the three different functional subsystems give different pictures on the phylogenetic relationship of organisms, reflecting the different extents of HGT in the different functional systems. In general, horizontal transfer is prevailing in genes for metabolism, but less in genes for information processing. Nevertheless, the major results of metabolic network-based phylogenetic trees are in good agreement with the tree based on 16S rRNA and genome trees, confirming the three domain classification and the close relationship between eukaryotes and archaea at the level of metabolic networks. These results strongly support the hypothesis that although HGT is widely distributed, it is nevertheless constrained by certain pre-existing metabolic organization principle(s) during the evolution. Further research is needed to identify the organization principle and constraints of metabolic network on HGT which have large impacts on understanding the evolution of life and in purposefully manipulating cellular metabolism.

Animals↗

Characterization in vitro and in vivo of the putative multigene 4-coumarate:CoA ligase network in Arabidopsis: syringyl lignin and sinapate/sinapyl alcohol derivative formation.

A recent in silico analysis revealed that the Arabidopsis genome has 14 genes annotated as putative 4-coumarate:CoA ligase isoforms or homologues. Of these, 11 were selected for detailed functional analysis in vitro, using all known possible phenylpropanoid pathway intermediates (p-coumaric, caffeic, ferulic, 5-hydroxyferulic and sinapic acids), as well as cinnamic acid. Of the 11 recombinant proteins so obtained, four were catalytically active in vitro, with fairly broad substrate specificities, confirming that the 4CL gene family in Arabidopsis has only four members. This finding is in agreement with our previous phylogenetic analyses, and again illustrates the need for comprehensive characterization of all putative 4CLs, rather than piecemeal analysis of selected gene members. All 11 proteins were expressed with a C-terminal His6-tag and functionally characterized, with one, At4CL1, expressed in native form for kinetic property comparisons. Of the 11 putative His6-tagged 4CLs, isoform At4CL1 best utilized p-coumaric, caffeic, ferulic and 5-hydroxyferulic acids as substrates, whereas At4CL2 readily transformed p-coumaric and caffeic acids into the corresponding CoA esters, while ferulic and 5-hydroxyferulic acids were converted quite poorly. At4CL3 also displayed broad substrate specificity efficiently converting p-coumaric, caffeic and ferulic acids into their CoA esters, whereas 5-hydroxyferulic acid was not as effectively utilized. By contrast, while At4CL5 is the only isoform capable of ligating sinapic acid, the two preferred substrates were 5-hydroxyferulic and caffeic acids. Indeed, both At4CL1 and At4CL5 most effectively utilized 5-hydroxyferulic acid with kenz approximately 10-fold higher than that for At4CL2 and At4CL3. The remaining seven 4CL-like homologues had no measurable catalytic activity (at approximately 100 microg protein concentrations), again bringing into sharp focus both the advantages to, and the limitations of, current database annotations, and the need to unambiguously demonstrate true enzyme function. Lastly, although At4CL5 is able to convert both 5-hydroxyferulic and sinapic acids into the corresponding CoA esters, the physiological significance of the latter observation in vitro was in question, i.e. particularly since other 4CL isoforms can effectively convert 5-hydroxyferulic acid into 5-hydroxyferuloyl CoA. Hence, homozygous lines containing T-DNA or enhancer trap inserts (knockouts) for 4cl5 were selected by screening, with Arabidopsis stem sections from each mutant line subjected to detailed analyses for both lignin monomeric compositions and contents, and sinapate/sinapyl alcohol derivative formation, at different stages of growth and development until maturation. The data so obtained revealed that this "knockout" had no significant effect on either lignin content or monomeric composition, or on the accumulation of sinapate/sinapyl alcohol derivatives. The results from the present study indicate that formation of syringyl lignins and sinapate/sinapyl alcohol derivatives result primarily from methylation of 5-hydroxyferuloyl CoA or derivatives thereof rather than sinapic acid ligation. That is, no specific physiological role for At4CL5 in direct sinapic acid CoA ligation could be identified. How the putative overlapping 4CL metabolic networks are in fact organized in planta at various stages of growth and development will be the subject of future inquiry.

Alcohols↗

Uncovering conserved patterns in bioactive peptides in Metazoa.

Bioactive (neuro)peptides play critical roles in regulating most biological processes in animals. Peptides belonging to the same family are characterized by a typical sequence pattern that is conserved among the family's peptide members. Such a conserved pattern or motif usually corresponds to the functionally important part of the biologically active peptide. In this paper, all known bioactive (neuro)peptides annotated in Swiss-Prot and TrEMBL protein databases are collected, and the pattern searching program Pratt is used to search these unaligned peptide sequences for conserved patterns. The obtained patterns are then refined by combining the information on amino acids at important functional sites collected from the literature. All the identified patterns are further tested by scanning them against Swiss-Prot and TrEMBL protein databases. The diagnostic power of each pattern is validated by the fact that any annotated protein from Swiss-Prot and TrEMBL that contains one of the established patterns, is indeed a known (neuro)peptide precursor. We discovered 155 novel peptide patterns in addition to the 56 established ones in the PROSITE database. All the patterns cover 110 peptide families. Fifty-five of these families are not characterized by the PROSITE signatures, and 12 are also not identified by other existing motif databases, such as Pfam and SMART. Using the newly identified peptide signatures as a search tool, we predicted 95 hypothetical proteins as putative peptide precursors.

Amino Acid Motifs↗

Gram-positive DsbE proteins function differently from Gram-negative DsbE homologs. A structure to function analysis of DsbE from Mycobacterium tuberculosis.

Mycobacterium tuberculosis, a Gram-positive bacterium, encodes a secreted Dsb-like protein annotated as Mtb DsbE (Rv2878c, also known as MPT53). Because Dsb proteins in Escherichia coli and other bacteria seem to catalyze proper folding during protein secretion and because folding of secreted proteins is thought to be coupled to disulfide oxidoreduction, the function of Mtb DsbE may be to ensure that secreted proteins are in their correctly folded states. We have determined the crystal structure of Mtb DsbE to 1.1 A resolution, which reveals a thioredoxin-like domain with a typical CXXC active site. These cysteines are in their reduced state. Biochemical characterization of Mtb DsbE reveals that this disulfide oxidoreductase is an oxidant, unlike Gram-negative bacteria DsbE proteins, which have been shown to be weak reductants. In addition, the pK(a) value of the active site, solvent-exposed cysteine is approximately 2 pH units lower than that of Gram-negative DsbE homologs. Finally, the reduced form of Mtb DsbE is more stable than the oxidized form, and Mtb DsbE is able to oxidatively fold hirudin. Structural and biochemical analysis implies that Mtb DsbE functions differently from Gram-negative DsbE homologs, and we discuss its possible functional role in the bacterium.

Amino Acid Sequence↗

Genome-Resolved Functional Profiling of Osteoporosis-Associated Gut Bacteria Highlights Putative Metabolic and Immunogenic Signatures of the Gut-Bone Axis.

The gut microbiota has emerged as a potential regulator of bone metabolism, but the genome-encoded functional repertoire of osteoporosis-associated gut bacteria remains insufficiently characterized. This study performed in silico functional profiling of gut bacterial taxa associated with osteoporosis, low bone mineral density, or comparator bone-related phenotypes. Twenty candidate taxa were selected from evidence in the human microbiome and represented by 26 curated bacterial reference genomes. Genome-wide annotations were used to map predicted gut-bone axis signatures, carbohydrate-active enzyme (CAZyme) repertoires, selected Kyoto Encyclopedia of Genes and Genomes pathways, and gutSMASH-predicted metabolic gene clusters. Functional burdens were normalized as hits per 1000 annotated proteins and integrated into metabolic, immunogenic, CAZyme, KEGG, and metabolic gene cluster profiles. Twelve predicted gut-bone axis signatures were identified, comprising 3337 primary candidate protein hits and a strict high-confidence subset of 2497 hits. Dominant signatures included vitamin B12/cobalamin metabolism, folate/one-carbon metabolism, peptidoglycan/cell-wall biosynthesis, and short-chain fatty acid-related functions. Dialister invisus, Dialister succinatiphilus, Megamonas funiformis, and Megamonas hypermegale showed the strongest normalized predicted gut-bone axis signal. These hypothesis-generating findings prioritize microbial metabolic and immunogenic features for future metagenomic, metabolomic, and experimental validation studies.

Osteoporosis↗

The tegument surface membranes of the human blood parasite Schistosoma mansoni: a proteomic analysis after differential extraction.

The blood fluke Schistosoma mansoni can live for years in the hepatic portal system of its human host and so must possess very effective mechanisms of immune evasion. The key to understanding how these operate lies in defining the molecular organisation of the exposed parasite surface. The adult worm is covered by a syncytial tegument, bounded externally by a plasma membrane and overlain by a laminate secretion, the membranocalyx. In order to determine the protein composition of this surface, the membranes were detached using a freeze/thaw technique and enriched by sucrose density gradient centrifugation. The resulting preparation was sequentially extracted with three reagents of increasing solubilising power. The extracts were separated by 2-DE and their protein constituents were identified by MS/MS, yielding predominantly cytosolic, cytoskeletal and membrane-associated proteins, respectively. After extraction, the final pellet containing membrane-spanning proteins was processed by liquid chromatographic techniques before MS. Transporters for sugars, amino acids, ions and other solutes were found together with membrane enzymes and proteins concerned with membrane structure. The proteins identified were categorised by their function and putative location on the basis of their homology with annotated proteins in other organisms.

Animals↗

Metagenomic analysis of microbial community dynamics in konjac rhizosphere during soft rot disease progression.

Amorphophallus konjac, the sole glucomannan-rich species in the Araceae family, faces significant yield and quality losses due to soft rot disease. Understanding the relationship between soil microbial communities and soft rot incidence is critical for sustainable konjac production. Metagenomic profiling was employed to systematically characterize the spatiotemporal dynamics of rhizosphere microbiomes during disease progression. Microbial alpha diversity (Chao1 index) exhibited a significant peak in the rhizosphere of diseased plants at the mature stage, contrasting with stable diversity patterns in healthy and latently infected groups, indicating dysbiosis-associated richness inflation during disease progression. Principal coordinate analysis (PCoA) revealed significant divergence in rhizosphere microbial structures between diseased and healthy/latently infected groups, with higher compositional variability observed in diseased samples. At the phylum level, Chloroflexi and Acidobacteria abundances in healthy mature plants exceeded those in diseased plants by 11.54% and 4.6%, respectively, while pathogenic Rhizopus arrhizus and Rhizopus microsporus were significantly enriched in diseased mature plants. Correlation analyses demonstrated predominantly negative associations between bacterial species and soil factors, contrasting with positive fungal correlations. KEGG pathway annotation identified carbohydrate metabolism and amino acid synthesis as core microbial functions in the konjac rhizosphere. Collectively, Chloroflexi and Acidobacteria were validated as putative biocontrol agents, while Rhizopus spp. emerged as key drivers of soft rot development. These findings provide mechanistic insights for designing microbiome-based biocontrol strategies to mitigate konjac soft rot, offering a sustainable alternative to conventional agrochemical reliance. KEY POINTS: • Diseased konjac microbial richness peaks; healthy plants enrich Chloroflexi/Acidobacteria. • Rhizopus pathogens drive soft rot; bacteria and fungi show opposing soil factor links. • Lays groundwork for microbiome approaches to cut agrochemicals in konjac rot control.

Rhizosphere↗

Combining transcriptome data with genomic and cDNA sequence alignments to make confident functional assignments for Aspergillus nidulans genes.

Whole genome sequencing of several filamentous ascomycetes is complete or in progress; these species, such as Aspergillus nidulans, are relatives of Saccharomyces cerevisiae. However, their genomes are much larger and their gene structure more complex, with genes often containing multiple introns. Automated annotation programs can quickly identify open reading frames for hypothetical genes, many of which will be conserved across large evolutionary distances, but further information is required to confirm functional assignments. We describe a comparative and functional genomics approach using sequence alignments and gene expression data to predict the function of Aspergillus nidulans genes. By highlighting examples of discrepancies between the automated genome annotation and cDNA or EST sequencing, we demonstrate that the greater complexity of gene structure in filamentous fungi demands independent data on gene expression and the gene sequence be used to make confident functional assignments.

Aspergillus nidulans↗

Predicting the solvent accessibility of transmembrane residues from protein sequence.

In this study, we propose a novel method to predict the solvent accessible surface areas of transmembrane residues. For both transmembrane alpha-helix and beta-barrel residues, the correlation coefficients between the predicted and observed accessible surface areas are around 0.65. On the basis of predicted accessible surface areas, residues exposed to the lipid environment or buried inside a protein can be identified by using certain cutoff thresholds. We have extensively examined our approach based on different definitions of accessible surface areas and a variety of sets of control parameters. Given that experimentally determining the structures of membrane proteins is very difficult and membrane proteins are actually abundant in nature, our approach is useful for theoretically modeling membrane protein tertiary structures, particularly for modeling the assembly of transmembrane domains. This approach can be used to annotate the membrane proteins in proteomes to provide extra structural and functional information.

Cell Membrane↗

Zinc through the three domains of life.

Zinc is one of the metal ions essential for life, as it is required for the proper functioning of a large number of proteins. Despite its importance, the annotation of zinc-binding proteins in gene banks or protein domain databases still has significant room for improvement. In the present work, we compiled a list of known zinc-binding protein domains and of known zinc-binding sequence motifs (zinc-binding patterns), and then used them jointly to analyze the proteome of 57 different organisms to obtain an overview of zinc usage by archaeal, bacterial, and eukaryotic organisms. Zinc-binding proteins are an abundant fraction of these proteomes, ranging between 4% and 10%. The number of zinc-binding proteins correlates linearly with the total number of proteins encoded by the genome of an organism, but the proportionality constant of Eukaryota (8.8%) is significantly higher than that observed in Bacteria and Archaea (from 5% to 6%). Most of this enrichment is due to the larger portfolio of regulatory proteins in Eukaryota.

Animals↗

A genomewide screen for components of the RNAi pathway in Drosophila cultured cells.

Posttranscriptional silencing by RNAi is initiated by dsRNAs that are processed into siRNAs that ultimately target homologous mRNAs for degradation. We used luciferase reporter constructs and a cultured cell-based assay to perform a genomewide screen for components of the RNAi pathway in Drosophila melanogaster. The screen identified seven genes that affect the RNAi response, five with previously described function (AGO2, Tis11, Hsc70-3, Hsc70-4, and hdc) and two annotated genes (CG17265 and CG10883).

Animals↗

TRILOGY: Discovery of sequence-structure patterns across diverse proteins.

We describe a new computer program, trilogy, for the automated discovery of sequence-structure patterns in proteins. trilogy implements a pattern discovery algorithm that begins with an exhaustive analysis of flexible three-residue patterns; a subset of these patterns are selected as seeds for an extension process in which longer patterns are identified. A key feature of the method is explicit treatment of both the sequence and structure components of these motifs: each trilogy pattern is a pair consisting of a sequence pattern and a structure pattern. Matches to both these component patterns are identified independently, allowing the program to assign a significance score to each sequence-structure pattern that assesses the degree of correlation between the corresponding sequence and structure motifs. trilogy identifies several thousand high-scoring patterns that occur across protein families. These include both previously identified and potentially novel motifs. We expect that these sequence-structure patterns will be useful in predicting protein structure from sequence, annotating newly determined protein structures, and identifying novel motifs of potential functional or structural significance. Further details on 7,768 significant patterns identified by trilogy can be found at http://theory.lcs.mit.edu/trilogy.

Algorithms↗

Identification of unstable transcripts in Arabidopsis by cDNA microarray analysis: rapid decay is associated with a group of touch- and specific clock-controlled genes.

mRNA degradation provides a powerful means for controlling gene expression during growth, development, and many physiological transitions in plants and other systems. Rates of decay help define the steady state levels to which transcripts accumulate in the cytoplasm and determine the speed with which these levels change in response to the appropriate signals. When fast responses are to be achieved, rapid decay of mRNAs is necessary. Accordingly, genes with unstable transcripts often encode proteins that play important regulatory roles. Although detailed studies have been carried out on individual genes with unstable transcripts, there is limited knowledge regarding their nature and associations from a genomic perspective, or the physiological significance of rapid mRNA turnover in intact organisms. To address these problems, we have applied cDNA microarray analysis to identify and characterize genes with unstable transcripts in Arabidopsis thaliana (AtGUTs). Our studies showed that at least 1% of the 11,521 clones represented on Arabidopsis Functional Genomics Consortium microarrays correspond to transcripts that are rapidly degraded, with estimated half-lives of less than 60 min. AtGUTs encode proteins that are predicted to participate in a broad range of cellular processes, with transcriptional functions being over-represented relative to the whole Arabidopsis genome annotation. Analysis of public microarray expression data for these genes argues that mRNA instability is of high significance during plant responses to mechanical stimulation and is associated with specific genes controlled by the circadian clock.

Arabidopsis↗

Additional gene ontology structure for improved biological reasoning.

MOTIVATION: The Gene Ontology (GO) is a widely used terminology for gene product characterization in, for example, interpretation of biology underlying microarray experiments. The current GO defines term relationships within each of the independent subontologies: molecular function, biological process and cellular component. However, it is evident that there also exist biological relationships between terms of different subontologies. Our aim was to connect the three subontologies to enable GO to cover more biological knowledge, enable a more consistent use of GO and provide new opportunities for biological reasoning. RESULTS: We propose a new structure, the Second Gene Ontology Layer, capturing biological relations not directly reflected in the present ontology structure. Given molecular functions, these paths identify biological processes where the molecular functions are involved and cellular components where they are active. The current Second Layer contains 6271 validated paths, covering 54% of the molecular functions of GO and can be used to render existing gene annotation sets more complete and consistent. Applying Second Layer paths to a set of 4223 human genes, increased biological process annotations by 24% compared to publicly available annotations and reproduced 30% of them. AVAILABILITY: The Second GO is publicly available through the GO Annotation Toolbox (GOAT.no): http://www.goat.no.

Computational Biology↗

Augur--a computational pipeline for whole genome microbial surface protein prediction and classification.

UNLABELLED: The analysis of protein function is a challenge and a major bottleneck towards well-annotated and analysed microbial genomes. In particular, bacterial surface proteins present an opportunity for pharmacological intervention and vaccine development. We present Augur, an automatic prediction pipeline that integrates major surface prediction algorithms and enables comparative analysis, classification and visualization for gram-positive bacteria on a genomic scale. AVAILABILITY: http://bioinfo.mikrobio.med.uni-giessen.de/augur

Algorithms↗

PANDORA: keyword-based analysis of protein sets by integration of annotation sources.

Recent advances in high-throughput methods and the application of computational tools for automatic classification of proteins have made it possible to carry out large-scale proteomic analyses. Biological analysis and interpretation of sets of proteins is a time-consuming undertaking carried out manually by experts. We have developed PANDORA (Protein ANnotation Diagram ORiented Analysis), a web-based tool that provides an automatic representation of the biological knowledge associated with any set of proteins. PANDORA uses a unique approach of keyword-based graphical analysis that focuses on detecting subsets of proteins that share unique biological properties and the intersections of such sets. PANDORA currently supports SwissProt keywords, NCBI Taxonomy, InterPro entries and the hierarchical classification terms from ENZYME, SCOP and GO databases. The integrated study of several annotation sources simultaneously allows a representation of biological relations of structure, function, cellular location, taxonomy, domains and motifs. PANDORA is also integrated into the ProtoNet system, thus allowing testing thousands of automatically generated clusters. We illustrate how PANDORA enhances the biological understanding of large, non-uniform sets of proteins originating from experimental and computational sources, without the need for prior biological knowledge on individual proteins.

Computational Biology↗