Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

Plant metabolomics: towards biological function and mechanism.

Metabolite profiling is a fast growing technology and is useful for phenotyping and diagnostic analyses of plants. It is also rapidly becoming a key tool in functional annotation of genes and in the comprehensive understanding of the cellular response to biological conditions. Metabolomics approaches have recently been used to assess the natural variance in metabolite content between individual plants, an approach with great potential for the improvement of the compositional quality of crops. Here, we assess the contribution of metabolite profiling to these areas.

Botany↗

Molecular analyses of disease pathogenesis: application of bovine microarrays.

The molecular analysis of disease pathogenesis in cattle has been limited by the lack of availability of tools to analyze both host and pathogen responses. These limitations are disappearing with the advent of methodologies such as microarrays that facilitate rapid characterization of global gene expression at the level of individual cells and tissues. The present review focuses on the use of microarray technologies to investigate the functional pathogenomics of infectious disease in cattle. We discuss a number of unique issues that must be addressed when designing both in vitro and in vivo model systems to analyze host responses to a specific pathogen. Furthermore, comparative functional genomic strategies are discussed that can be used to address questions regarding host responses that are either common to a variety of pathogens or unique to individual pathogens. These strategies can also be applied to investigations of cell signaling pathways and the analyses of innate immune responses. Microarray analyses of both host and pathogen responses hold substantial promise for the generation of databases that can be used in the future to address a wide variety of questions. A critical component limiting these comparative analyses will be the quality of the databases and the complete functional annotation of the bovine genome. These limitations are discussed with an indication of future developments that will accelerate the validation of data generated when completing a molecular characterization of disease pathogenesis in cattle.

Animals↗

A conserved distal-tail helical extension defines a tailspike attachment architecture in Gram-negative siphophages.

Rapid growth of bacteriophage genome collections has outpaced functional annotation of tail-tip proteins, limiting comparative analysis of host-recognition structures. Starting from a shared distal-tail gene organization in the Salmonella phages 9NA and Jersey, I developed a morphogenetic bioinformatic framework integrating gene synteny, sequence comparison, profile hidden Markov model (HMM) screening, structural evidence, structure-aware searching, and AlphaFold modeling. Comparison with the experimentally characterized lambda and Sf11 tail assemblies identified a predominantly alpha-helical C-terminal extension of the distal-tail (DT) protein associated with tailspike attachment, termed the distal-tail helical extension (DT-helix). Screening 541,986 proteins from 5167 complete NCBI RefSeq tailed-phage genomes, followed by evidence-based evaluation of sequence, genomic context, and structural architecture, identified 165 curated DT-helical-extension-associated phages. Their DT proteins segregated into six sequence groups. In the four principal multi-member groups, cognate tailspikes showed group-specific conservation in proximal N-terminal regions but substantially greater downstream diversity, consistent with sequence constraint at the DT-tailspike attachment boundary. A complementary ProstT5/Foldseek search supported the established groups but revealed no convincing additional highly divergent family. Together with the experimentally characterized Sf11 attachment interface, these findings define a recurrent morphogenetic architecture linking conserved distal-tail scaffolds to more variable receptor-binding proteins across siphophages infecting Gram-negative bacteria. Although universal exchangeability is not established, the identified scaffold-receptor-binding boundaries provide a framework for molecular characterization and rational phage engineering. Accession-level information for the 165 curated phages is available through PhageTailDB.

Viral Tail Proteins↗

Transposable element-driven expansion of enhancer RNA repertoires underlies regulatory innovation and polyploid adaptation in cereal crops.

Cereal genomes have undergone repeated polyploidization and transposable element (TE) proliferation, collectively generating complex regulatory landscapes. However, the evolutionary trajectories and functional implications of these landscapes remain largely unexplored. Using chromatin-bound RNA sequencing across seven cereal species, we systematically mapped 45,952 regulatory element transcripts (RETs), including 32,867 distal RETs corresponding to enhancer RNAs (eRNAs). Our analysis revealed that 56% of lineage-specific eRNAs originated from TE expansions, indicating that TEs serve as major reservoirs of species-specific regulatory innovation in cereals. Notably, we identified remarkable conservation in defense-related functions, root-specific expression, and TE-derived origins of eRNAs across both ancient and recent evolutionary layers of Triticeae, suggesting recurrent recruitment of TE-derived, root-associated regulatory elements throughout Triticeae evolution. Furthermore, we found that young eRNA pairs in hexaploid wheat with high sequence similarity, many originating from RLG_famc8.3 and DTC_famc4.3, exhibited pronounced root specificity and coordinated expression, suggesting targeted amplification and refinement of successful ancestral regulatory strategies established after Triticeae divergence. To facilitate community access, we developed Cereal-eRNAdb (http://bioinfo.cemps.ac.cn/Cereal-eRNAdb/), a comprehensive database integrating 69,426 eRNAs with functional annotations across 296 samples. Our findings suggest that TE-mediated innovation of root-specific eRNAs may contribute to Triticeae adaptation and provide a foundational resource for exploiting regulatory variation in cereal crop breeding.

Enhancer RNAs↗

IRIS: a database surveying known human immune system genes.

We have compiled an online database of known human defense genes: the Immunogenetic Related Information Source (IRIS). As of October 1, 2004, there are 1562 immune genes recorded in IRIS, representing 7% of the human genome. This resource contains searchable information including chromosomal location, sequence data, and a curated functional annotation for each entry. We used IRIS as a basis for analyzing the composition and characteristics of the immune genome, such as gene clustering, polymorphism, and relationship to disease. High protein sequence similarity correlated inversely with distance between immune genes, consistent with clustering of duplicated loci. We also found that, even though some immune genes exhibit high levels of polymorphism, such as MHC class I, the range of levels of polymorphism in immune genes is similar to that of nonimmune genes. Approximately 20% of immune genes have a known disease association. IRIS is available online at .

Databases, Genetic↗

ChickGCE: a novel germ cell EST database for studying the early developmental stage in chickens.

We established a database to study germ cells during the early developmental stage in the chicken. The ChickGCE database provides integrated expressed sequence tag (EST) data from chicken testis, ovary, embryonic gonads, and primordial germ cells. We gathered data on 10,294 ESTs from approximately 1000 embryonic gonads, and we experimentally determined 10,851 ESTs from primordial germ cells purified from 7955 embryonic gonads by magnetically activated cell sorting. The EST testis and ovary datasets were retrieved from the public database of The Institute for Genomic Research (TIGR). The EST data were clustered and assembled into unique sequences, contigs, and singletons. The ChickGCE database provides functional annotation, identification, and putative embryonic germ-cell-specific novel transcripts based on the Gene Ontology database, as well as statistical analyses of expression patterns and pair-wise comparisons of two types of tissue- and germ-cell-specific alternative splicing events in the chicken. The new database is accessible online and queries can be answered using several search options, including tissue database searches, keywords, clone IDs, expected values, and BLAST search scores.

Animals↗

Role of context in the relationship between form and function: structural plasticity of some PROSITE patterns.

True positive hits of PROSITE sequence pattern are expected to have a characteristic three-dimensional structure. The combined sequence-structure attributes of PROSITE patterns can be used for function prediction of an uncharacterized protein with known primary and 3D structure, a situation that might arise in structural genomics projects. We have found specific examples of true hits of PROSITE patterns displaying structural plasticity by assuming significantly different local conformation, depending upon the context. Our work highlights the importance of taking into account all the known distinct conformations of PROSITE patterns, while creating a sensitive 3D template for the pattern, for use in functional annotation.

Amino Acid Motifs↗

A new family of CoA-transferases.

CoA-transferases are found in organisms from all lines of descent. Most of these enzymes belong to two well-known enzyme families, but recent work on unusual biochemical pathways of anaerobic bacteria has revealed the existence of a third family of CoA-transferases. The members of this enzyme family differ in sequence and reaction mechanism from CoA-transferases of the other families. Currently known enzymes of the new family are a formyl-CoA: oxalate CoA-transferase, a succinyl-CoA: (R)-benzylsuccinate CoA-transferase, an (E)-cinnamoyl-CoA: (R)-phenyllactate CoA-transferase, and a butyrobetainyl-CoA: (R)-carnitine CoA-transferase. In addition, a large number of proteins of unknown or differently annotated function from Bacteria, Archaea and Eukarya apparently belong to this enzyme family. Properties and reaction mechanisms of the CoA-transferases of family III are described and compared to those of the previously known CoA-transferases.

Bacteria, Anaerobic↗

A profile of differentially expressed genes in primary colorectal cancer using suppression subtractive hybridization.

As a step towards understanding the complex differences between normal cells and cancer cells, we have used suppression subtractive hybridization (SSH) to generate a profile of genes overexpressed in primary colorectal cancer (CRC). From a 35¿ omitted¿000 clone SSH-cDNA repertoire, we have screened 400 random clones by reverse Northern blotting, of which 45 clones were scored as overexpressed in tumor compared to matched normal mucosa. Sequencing showed 37 different genes and of these, 16 genes corresponded to known genes in the public databases. Twelve genes, including Smad5 and Fls353, have previously been shown to be overexpressed in CRC. A series of known genes which have not previously been reported to be overexpressed in cancer were also recovered: Hsc70, PBEF, ribophorin II and Ese-3B. The remaining 21 genes have as yet no functional annotation. These results show that SSH in conjunction with high throughput screening provides a very efficient means to produce a broad profile of genes differentially expressed in cancer. Some of the genes identified may provide novel points of therapeutic intervention.

Adenocarcinoma↗

An accurate, sensitive, and scalable method to identify functional sites in protein structures.

Functional sites determine the activity and interactions of proteins and as such constitute the targets of most drugs. However, the exponential growth of sequence and structure data far exceeds the ability of experimental techniques to identify their locations and key amino acids. To fill this gap we developed a computational Evolutionary Trace method that ranks the evolutionary importance of amino acids in protein sequences. Studies show that the best-ranked residues form fewer and larger structural clusters than expected by chance and overlap with functional sites, but until now the significance of this overlap has remained qualitative. Here, we use 86 diverse protein structures, including 20 determined by the structural genomics initiative, to show that this overlap is a recurrent and statistically significant feature. An automated ET correctly identifies seven of ten functional sites by the least favorable statistical measure, and nine of ten by the most favorable one. These results quantitatively demonstrate that a large fraction of functional sites in the proteome may be accurately identified from sequence and structure. This should help focus structure-function studies, rational drug design, protein engineering, and functional annotation to the relevant regions of a protein.

Amino Acid Motifs↗

Engineering embryonic stem cells with recombinase systems.

The combined use of site-specific recombination and gene targeting or trapping in embryonic stem cells (ESCs) has resulted in the emergence of technologies that enable the induction of mouse mutations in a prespecified temporal and spatially restricted manner. Their large-scale implementation by several international mouse mutagenesis programs will lead to the assembly of a library of ES cell lines harboring conditional mutations in every single gene of the mouse genome. In anticipation of this unprecedented resource, this chapter will focus on site-specific recombination strategies and issues pertinent to ESCs and mice. The upcoming ESC resource and the increasing sophistication of site-specific recombination technologies will greatly assist the functional annotation of the human genome and the animal modeling of human disease.

Animals↗

Gene networks: how to put the function in genomics.

An increasingly popular model of regulation is to represent networks of genes as if they directly affect each other. Although such gene networks are phenomenological because they do not explicitly represent the proteins and metabolites that mediate cell interactions, they are a logical way of describing phenomena observed with transcription profiling, such as those that occur with popular microarray technology. The ability to create gene networks from experimental data and use them to reason about their dynamics and design principles will increase our understanding of cellular function. We propose that gene networks are also a good way to describe function unequivocally, and that they could be used for genome functional annotation. Here, we review some of the concepts and methods associated with gene networks, with emphasis on their construction based on experimental data.

Animals↗

Identification of Gal80p-interacting proteins by Saccharomyces cerevisiae whole genome phage display.

Networks of interacting proteins and protein interaction maps can help in functional annotation in genome analysis projects. We present the application of genomic phage display as a tool to identify interacting proteins in Saccharomyces cerevisiae. We have developed a large phagemid display library (approximately 7.7x10(7) independent clones) of sheared S. cerevisiae genomic DNA (12.1 Mbp genome size) fused to gene III (lacking the N1 domain) of the filamentous phage M13. Baits tagged with an N-terminal E-tag and a C-terminal His(6)-tag are prepared in a novel Escherichia coli expression system. Using E-Gal80-His(6) as bait, biopanning of the library resulted in the isolation of two different clones containing fragments of the known interacting partner Gal4p. In addition, three new ligands (Ubr1p, YCL045c and Prp8p) with potential physiological relevance were isolated. Interactions were confirmed by ELISA. These results demonstrate the accessibility of the S. cerevisiae genome to display technology for protein-protein interaction screening.

Amino Acid Sequence↗

Identification of 9 novel transcripts and two RGSL genes within the hereditary prostate cancer region (HPC1) at 1q25.

We applied a systematic bioinformatics approach, followed by careful manual inspection and experimental validation to identify additional expressed sequences located at the Hereditary Prostate Cancer Region (HPC1) between D1S2818 and D1S1642 on chromosome 1q25. All transcripts already described for the 1q25 region were identified and we were able to define 11 additional expressed sequences within this region (three full-length cDNA clone sequences and eight ESTs), increasing the total number of gene count in this region by 38%. Five out of the 11 expressed sequences identified were shown to be expressed in prostate tissue and thus represent novel disease gene candidates for the HPC1 region. Here, we report a detailed characterization of these five novel disease gene candidates, their expression pattern in various tissues, their genomic organization and functional annotation. Two candidates (RGSL1 and RGSL2) correspond to novel members of the RGS family, which is involved in the regulation of G-protein signaling. RGSL1 and RGLS2 expression was detected by real-time polymerase chain reaction in normal prostate tissue, but could not be detected in prostate tumor cell lines, suggesting they might have a role in prostate cancer.

Chromosome Mapping↗

Mass spectrometric analysis of the editosome and other multiprotein complexes in Trypanosoma brucei.

The composition of the editosome, a multi-protein complex that catalyzes uridine insertion and deletion RNA editing to produce mature mitochondrial mRNAs in trypanosomes, was analyzed by mass spectrometry. The editosomes were isolated by column chromatography, glycerol gradient sedimentation, and monoclonal antibody affinity purifications. At least 16 proteins form the catalytic core of the editosome, and additional associated proteins were identified. Analyses of mitochondrial fractions identified several non-editosome proteins and multi-protein complexes. These studies contribute to the functional annotation of T. brucei genome.

Amino Acid Sequence↗

Genome-wide ENU mutagenesis to reveal immune regulators.

A complete list of molecular components for immune system function is now available with the completion of the human and mouse genome sequences. However, identification and functional annotation of genes involved in immunological processes require a discovery methodology that can efficiently and broadly analyze the complex interplay of these components in vivo. Our recent experience indicates that genome-wide chemical mutagenesis in the mouse is an extremely powerful methodology for the identification of genes required for complex immunological processes.

Animals↗

Protein-interaction networks: from experiments to analysis.

Functional proteomics approaches aim to characterize comprehensively the function of gene products, and provide a first-level understanding of cellular mechanisms. Here, we review recent techniques for the construction and prediction of large-scale protein-interaction networks, with a particular emphasis on computational processing steps and comparative assessment of the reliability and completeness of the various approaches. We also discuss the use of protein-interaction network information in functional annotation and in the generation of higher-level biological hypotheses on pathways.

Computational Biology↗

Predicting the subcellular localization of human proteins using machine learning and exploratory data analysis.

Identifying the subcellular localization of proteins is particularly helpful in the functional annotation of gene products. In this study, we use Machine Learning and Exploratory Data Analysis (EDA) techniques to examine and characterize amino acid sequences of human proteins localized in nine cellular compartments. A dataset of 3,749 protein sequences representing human proteins was extracted from the SWISS-PROT database. Feature vectors were created to capture specific amino acid sequence characteristics. Relative to a Support Vector Machine, a Multi-layer Perceptron, and a Naive Bayes classifier, the C4.5 Decision Tree algorithm was the most consistent performer across all nine compartments in reliably predicting the subcellular localization of proteins based on their amino acid sequences (average Precision=0.88; average Sensitivity=0.86). Furthermore, EDA graphics characterized essential features of proteins in each compartment. As examples, proteins localized on the plasma membrane had higher proportions of hydrophobic amino acids; cytoplasmic proteins had higher proportions of neutral amino acids; and mitochondrial proteins had higher proportions of neutral amino acids and lower proportions of polar amino acids. These data showed that the C4.5 classifier and EDA tools can be effective for characterizing and predicting the subcellular localization of human proteins based on their amino acid sequences.

Algorithms↗