Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,153 records · Page 64Linked to original sources

Gene expression profiling identifies molecular subgroups among nodal peripheral T-cell lymphomas.

The classification of peripheral T-cell lymphomas (PTCL) is still a matter of debate. To establish a molecular classification of PTCL, we analysed 59 primary nodal T-cell lymphomas using cDNA microarrays, including 56 PTCL and three T-lymphoblastic lymphoma (T-LBL). The expression profiles could discriminate angioimmunoblastic lymphoma, anaplastic large-cell lymphoma and T-LBL. In contrast, cases belonging to the broad category of 'PTCL, unspecified' (PTCL-U) did not share a single molecular profile. Using a multiclass predictor, we could separate PTCL-U into three molecular subgroups called U1, U2 and U3. The U1 gene expression signature included genes known to be associated with poor outcome in other tumors, such as CCND2. The U2 subgroup was associated with overexpression of genes involved in T-cell activation and apoptosis, including NFKB1 and BCL-2. The U3 subgroup was mainly defined by overexpression of genes involved in the IFN/JAK/STAT pathway. It comprised a majority of histiocyte-rich PTCL samples. Gene Ontology annotations revealed different functional profile for each subgroup. These results suggest the existence of distinct subtypes of PTCL-U with specific molecular profiles, and thus provide a basis to improve their classification and to develop new therapeutic targets.

Gene Expression Profiling↗

Microarray analyses of gene expression during chondrocyte differentiation identifies novel regulators of hypertrophy.

Ordered chondrocyte differentiation and maturation is required for normal skeletal development, but the intracellular pathways regulating this process remain largely unclear. We used Affymetrix microarrays to examine temporal gene expression patterns during chondrogenic differentiation in a mouse micromass culture system. Robust normalization of the data identified 3300 differentially expressed probe sets, which corresponds to 1772, 481, and 249 probe sets exhibiting minimum 2-, 5-, and 10-fold changes over the time period, respectively. GeneOntology annotations for molecular function show changes in the expression of molecules involved in transcriptional regulation and signal transduction among others. The expression of identified markers was confirmed by RT-PCR, and cluster analysis revealed groups of coexpressed transcripts. One gene that was up-regulated at later stages of chondrocyte differentiation was Rgs2. Overexpression of Rgs2 in the chondrogenic cell line ATDC5 resulted in accelerated hypertrophic differentiation, thus providing functional validation of microarray data. Collectively, these analyses provide novel information on the temporal expression of molecules regulating endochondral bone development.

Animals↗

A tool to assist the study of specific features at protein binding sites.

The Protein Data Bank contains a large amount of proteins that have been solved with small ligands bound to them. This constitutes a rich source of information for the study of the specific requirements of protein sites to bind small molecules with a favorable free energy. The specific atomic composition and three-dimensional geometric restraints of protein binding sites for different ligands could be easily obtained from there. The development of accurate binding site descriptors in proteins constitutes a valuable tool to assist in the large-scale prediction and annotation of protein function in whole genomes. In this work, an integrated database containing some processed and calculated protein/ligand information is described. It is expected that this database will constitute a useful tool for people working in the prediction of protein function from its structure. The database is accessible from the Internet through a web server located at: http://protein.bio.puc.cl

Binding Sites↗

Bio-support vector machines for computational proteomics.

MOTIVATION: One of the most important issues in computational proteomics is to produce a prediction model for the classification or annotation of biological function of novel protein sequences. In order to improve the prediction accuracy, much attention has been paid to the improvement of the performance of the algorithms used, few is for solving the fundamental issue, namely, amino acid encoding as most existing pattern recognition algorithms are unable to recognize amino acids in protein sequences. Importantly, the most commonly used amino acid encoding method has the flaw that leads to large computational cost and recognition bias. RESULTS: By replacing kernel functions of support vector machines (SVMs) with amino acid similarity measurement matrices, we have modified SVMs, a new type of pattern recognition algorithm for analysing protein sequences, particularly for proteolytic cleavage site prediction. We refer to the modified SVMs as bio-support vector machine. When applied to the prediction of HIV protease cleavage sites, the new method has shown a remarkable advantage in reducing the model complexity and enhancing the model robustness.

Algorithms↗

Integrative Array Analyzer: a software package for analysis of cross-platform and cross-species microarray data.

The rapid accumulation of microarray data translates into an urgent need for tools to perform integrative microarray analysis. Integrative Array Analyzer is a comprehensive analysis and visualization software toolkit, which aims to facilitate the reuse of the large amount of cross-platform and cross-species microarray data. It is composed of the data preprocess module, the co-expression analysis module, the differential expression analysis module, the functional and transcriptional annotation module and the graph visualization module.

Algorithms↗

Integration of genome mining and HiTES reveals secondary metabolic potential in marine-derived Aspergillus sp. WHUF0304.

AIMS: Marine-derived Aspergillus species are prolific producers of bioactive secondary metabolites, yet the majority of their biosynthetic gene clusters (BGCs) remain silent. This study aimed to integrate genome mining with high-throughput elicitor screening (HiTES) to unlock the metabolic potential of Aspergillus sp. WHUF0304 and identify elicitors that promote the accumulation of previously undetected metabolites. METHODS AND RESULTS: A high-quality genome of Aspergillus sp. WHUF0304 was assembled and annotated using multiple functional databases, revealing substantial secondary metabolic potential. antiSMASH analysis identified diverse BGCs, including NRPS/indole-related clusters potentially associated with indole diketopiperazine biosynthesis. A HiTES-inspired elicitor screening strategy was then applied to evaluate 42 small molecules for their ability to alter the metabolite profile of this strain. Among the tested elicitors, fluconazole was identified as the optimal inducer, triggering the production of several indole diketopiperazine-related differential metabolites. Subsequent activity-guided isolation led to the identification of a bioactive indole diketopiperazine dimer, cristatumin E, which exhibited antibacterial activity against Escherichia coli and Bacillus subtilis with minimum inhibitory concentrations (MICs) of 32 µg mL-1 and 256 µg mL-1, respectively. CONCLUSIONS: These findings demonstrate that integrating genomic and functional approaches effectively activates silent BGCs in marine fungi. The fluconazole-associated accumulation and subsequent isolation of cristatumin E, a bioactive indole diketopiperazine dimer, highlight the potential of elicitor-mediated activation to expand the detectable metabolite profile of Aspergillus sp. WHUF0304.

Aspergillus↗

The prostate expression database (PEDB): status and enhancements in 2000.

The Prostate Expression Database (PEDB) is an online resource designed to access and analyze gene expression information derived from the human prostate. PEDB archives >55 000 expressed sequence tags (ESTs) from 43 cDNA libraries in a curated relational database that provides detailed library information including tissue source, library construction methods, sequence diversity and sequence abundance. The differential expression of each EST species can be viewed across all libraries using a Virtual Expression Analysis Tool (VEAT), a graphical user interface written in Java for intra- and inter-library species comparisons. Recent enhancements to PEDB include: (i) the functional categorization of annotated EST assemblies using a classification scheme developed at The Institute for Genome Research; (ii) catalogs of expressed genes in specific prostate tissue sources designated as transcriptomes; and (iii) the addition of prostate proteome information derived from two-dimensional electrophoreses and mass spectrometry of prostate cancer cell lines. PEDB may be accessed via the WWW at http://www.mbt.washington.edu/PEDB/

Databases, Factual↗

The COG database: a tool for genome-scale analysis of protein functions and evolution.

Rational classification of proteins encoded in sequenced genomes is critical for making the genome sequences maximally useful for functional and evolutionary studies. The database of Clusters of Orthologous Groups of proteins (COGs) is an attempt on a phylogenetic classification of the proteins encoded in 21 complete genomes of bacteria, archaea and eukaryotes (http://www. ncbi.nlm. nih.gov/COG). The COGs were constructed by applying the criterion of consistency of genome-specific best hits to the results of an exhaustive comparison of all protein sequences from these genomes. The database comprises 2091 COGs that include 56-83% of the gene products from each of the complete bacterial and archaeal genomes and approximately 35% of those from the yeast Saccharomyces cerevisiae genome. The COG database is accompanied by the COGNITOR program that is used to fit new proteins into the COGs and can be applied to functional and phylogenetic annotation of newly sequenced genomes.

Database Management Systems↗

The SBASE domain sequence library, release 10: domain architecture prediction.

SBASE (http://www.icgeb.trieste.it/sbase) is an on-line collection of protein domain sequences and related computational tools designed to facilitate detection of domain homologies based on simple database search. The 10th 'jubilee release' of the SBASE library of protein domain sequences contains 1 052 904 protein sequence segments annotated by structure, function, ligand-binding or cellular topology, clustered into over 6000 domain groups. Domain identification and functional prediction are based on a comparison of BLAST search outputs with a knowledge base of biologically significant similarities extracted from known domain groups. The knowledge base is generated automatically for each domain group from the comparison of within-group ('self') and out-of-group ('non-self') similarities. This is a memory-based approach wherein group-specific similarity functions are automatically learned from the database.

Animals↗

ArchDB: automated protein loop classification as a tool for structural genomics.

The annotation of protein function has become a crucial problem with the advent of sequence and structural genomics initiatives. A large body of evidence suggests that protein structural information is frequently encoded in local sequences, and that folds are mainly made up of a number of simple local units of super-secondary structural motifs, consisting of a few secondary structures and their connecting loops. Moreover, protein loops play an important role in protein function. Here we present ArchDB, a classification database of structural motifs, consisting of one loop plus its bracing secondary structures. ArchDB currently contains 12,665 super-secondary elements classified into 1496 motif subclasses. The database provides an easy way to retrieve functional information from protein structures sharing a common motif, to search motifs found in a given SCOP family, superfamily or fold, or to search by keywords on proteins with classified loops. The ArchDB database of loops is located at http://sbi.imim.es/archdb.

Amino Acid Motifs↗

ECR Browser: a tool for visualizing and accessing data from comparisons of multiple vertebrate genomes.

With an increasing number of vertebrate genomes being sequenced in draft or finished form, unique opportunities for decoding the language of DNA sequence through comparative genome alignments have arisen. However, novel tools and strategies are required to accommodate this large volume of genomic information and to facilitate the transfer of predictions generated by comparative sequence alignment to researchers focused on experimental annotation of genome function. Here, we present the ECR Browser, a tool that provides easy and dynamic access to whole genome alignments of human, mouse, rat and fish sequences. This web-based tool (http://ecrbrowser.dcode.org) provides the starting point for discovery of novel genes, identification of distant gene regulatory elements and prediction of transcription factor binding sites. The genome alignment portal of the ECR Browser also permits fast and automated alignments of any user-submitted sequence to the genome of choice. The interconnection of the ECR Browser with other DNA sequence analysis tools creates a unique portal for studying and exploring vertebrate genomes.

Animals↗

The SBASE domain sequence resource, release 12: prediction of protein domain-architecture using support vector machines.

SBASE (http://www.icgeb.trieste.it/sbase) is an online resource designed to facilitate the detection of domain homologies based on sequence database search. The present release of the SBASE A library of protein domain sequences contains 972,397 protein sequence segments annotated by structure, function, ligand-binding or cellular topology, clustered into 8547 domain groups. SBASE B contains 169,916 domain sequences clustered into 2526 less well-characterized groups. Domain prediction is based on an evaluation of database search results in comparison with a 'similarity network' of inter-sequence similarity scores, using support vector machines trained on similarity search results of known domains.

Artificial Intelligence↗

GeneTrees: a phylogenomics resource for prokaryotes.

The GeneTrees phylogenomics system pursues comparative genomic analyses from the perspective of gene phylogenies for individual genes. The GeneTrees project has the goal of providing detailed evolutionary models for all protein-coding gene components of the fully sequenced genomes. Currently, a database of alignments and trees for all protein sequences for 325 fully sequenced and annotated prokaryote genomes is available. The prokaryote database contains 890,000 protein sequences organized into over 100,000 alignments, each described by a phylogenetic tree. An original homology group discovery tool assembles sets of related proteins from all versus all pairwise alignments. Multiple alignments for each homology group are stored and subjected to phylogenetic tree inference. A graphical web interface provides visual exploration of the GeneTrees database. Homology groups can be queried by sequence identifiers or annotation terms. Genomes can be browsed visually on a gene map of each chromosome or plasmid. Phylogenetic trees with support values are displayed in conjunction with the associated sequence alignment. A variety of classes of information can be selected to label the tree tips to aid in visual evaluation of annotation and gene function. This web interface is available at http://genetrees.vbi.vt.edu.

Bacterial Proteins↗

On the spatial disposition of the fifth transmembrane helix and the structural integrity of the transmembrane binding site in the opioid and ORL1 G protein-coupled receptor family.

Evidence from statistical cluster analyses of a multiple sequence alignment of G protein-coupled receptor seven-helix folds supports the existence of structurally conserved transmembrane (TM) ligand binding sites in the opioid/opioid receptor-like (ORL1) and amine receptor families. Based on the expectation that functionally conserved regions in homologous proteins will display locally higher levels of sequence identity compared with global sequence similarities that pertain to the overall fold, this approach may have wider applications in functional genomics to annotate sequence data. Binding sites in models of the kappa-opioid receptor seven-helix bundle built from the rhodopsin templates of Baldwin et al. (1997) [J. Mol. Biol., 272, 144-164] and Herzyk and Hubbard (1998) [J. Mol. Biol., 281, 742-751] are compared. The Herzyk and Hubbard template is found to be in better accord with experimental studies of amine, opioid and rhodopsin receptors owing to the reduced physical separation of the extracellular parts of TM helices V and VI and differences in the rotational orientation of the N-terminal of helix V that reveal side chain accessibilities in the Baldwin et al. structure to be out of phase with relative alkylation rates of engineered cysteine residues in the TM binding site of the alpha(2A)-adrenergic receptor. TM helix V in the Baldwin et al. template has been remodelled with a different proline kink to satisfy experimental constraints. A recent proposal that rotation of helix V is associated with receptor activation is critically discussed.

Binding Sites↗

Computing prokaryotic gene ubiquity: rescuing the core from extinction.

The genomic core concept has found several uses in comparative and evolutionary genomics. Defined as the set of all genes common to (ubiquitous among) all genomes in a phylogenetically coherent group, core size decreases as the number and phylogenetic diversity of the relevant group increases. Here, we focus on methods for defining the size and composition of the core of all genes shared by sequenced genomes of prokaryotes (Bacteria and Archaea). There are few (almost certainly less than 50) genes shared by all of the 147 genomes compared, surely insufficient to conduct all essential functions. Sequencing and annotation errors are responsible for the apparent absence of some genes, while very limited but genuine disappearances (from just one or a few genomes) can account for several others. Core size will continue to decrease as more genome sequences appear, unless the requirement for ubiquity is relaxed. Such relaxation seems consistent with any reasonable biological purpose for seeking a core, but it renders the problem of definition more problematic. We propose an alternative approach (the phylogenetically balanced core), which preserves some of the biological utility of the core concept. Cores, however delimited, preferentially contain informational rather than operational genes; we present a new hypothesis for why this might be so.

Conserved Sequence↗

Patterns of Genomic Divergence and Introgression in Two Primulina Hybrid Zones.

Hybrid zones have long been promoted as natural laboratories for understanding the mechanisms of speciation. Multiple or replicated hybrid zones are particularly informative, as they allow for assessing the consistency of genomic divergence and introgression across different environmental contexts and demographic histories, thereby improving our understanding of the factors that drive or hinder speciation on a broader scale. Here, using whole-genome resequencing data, we compare the patterns of genomic divergence and introgression in two Primulina hybrid zones. We found that genomic divergence in both hybrid zones is largely shaped by neutral processes, with only a few genomic regions showing signatures of balancing or lineage-specific selection. Genomic cline analyses identified numerous SNPs that showed significantly steeper clines and biased centres than the genome-wide expectation in both hybrid zones, consistent with the existence of reproductive barriers. Within regions of restricted gene flow, we identified 21 genes shared between the two hybrid zones. Annotation of gene function revealed that several genes are involved in reproductive processes. In addition, many zone-specific outlier loci were linked to genes associated with pollen and flower development, suggesting that these barriers may contribute to reproductive isolation under localised ecological conditions. Overall, these findings suggest that while certain reproductive barriers remain consistent across independent hybrid zones, others may be contingent on local environmental contexts. Our results demonstrate that both general and zone-specific mechanisms contribute to reproductive isolation in Primulina, providing empirical evidence that some genomic barriers recur across independent hybrid zones while others arise through localised adaptation.

Lamiales↗

Genes that enhance the ecological fitness of Shewanella oneidensis MR-1 in sediments reveal the value of antibiotic resistance.

Environmental bacteria persist in various habitats, yet little is known about the genes that contribute to growth and survival in their respective ecological niches. Signature-tagged mutagenesis (STM) of Shewanella oneidensis MR-1 coupled with a screen involving incubations of mutant strains in anoxic aquifer sediments allowed us to identify 47 genes that enhance fitness in sediments. Gene functions inferred from annotations provide us with insight into physiological and ecological processes that environmental bacteria use while growing in sediment ecosystems. Identification of the mexF gene and other potential membrane efflux components by STM demonstrated that homologues of multidrug resistance genes present in pathogens are required for sediment fitness of nonpathogenic bacteria. Further studies with a mexF deletion mutant demonstrated that the multidrug resistance pump encoded by mexF is required for resistance to antibiotics, including chloramphenicol and tetracycline. Chloramphenicol-adapted cultures exhibited mutations in the gene encoding a TetR family regulatory protein, indicating a role for this protein in regulating expression of the mexEF operon. The relative importance of mexF for sediment fitness suggests that antibiotic efflux may be a required process for bacteria living in sediment systems.

Anti-Bacterial Agents↗

A new pathway for salvaging the coenzyme B12 precursor cobinamide in archaea requires cobinamide-phosphate synthase (CbiB) enzyme activity.

The ability of archaea to salvage cobinamide has been under question because archaeal genomes lack orthologs to the bacterial nucleoside triphosphate:5'-deoxycobinamide kinase enzyme (cobU in Salmonella enterica). The latter activity is required for cobinamide salvaging in bacteria. This paper reports evidence that archaea salvage cobinamide from the environment by using a pathway different from the one used by bacteria. These studies demanded the functional characterization of two genes whose putative function had been annotated based solely on their homology to the bacterial genes encoding adenosylcobyric acid and adenosylcobinamide-phosphate synthases (cbiP and cbiB, respectively) of S. enterica. A cbiP mutant strain of the archaeon Halobacterium sp. strain NRC-1 was auxotrophic for adenosylcobyric acid, a known intermediate of the de novo cobamide biosynthesis pathway, but efficiently salvaged cobinamide from the environment, suggesting the existence of a salvaging pathway in this archaeon. A cbiB mutant strain of Halobacterium was auxotrophic for adenosylcobinamide-GDP, a known de novo intermediate, and did not salvage cobinamide. The results of the nutritional analyses of the cbiP and cbiB mutants suggested that the entry point for cobinamide salvaging is adenosylcobyric acid. The data are consistent with a salvaging pathway for cobinamide in which an amidohydrolase enzyme cleaves off the aminopropanol moiety of adenosylcobinamide to yield adenosylcobyric acid, which is converted by the adenosylcobinamide-phosphate synthase enzyme to adenosylcobinamide-phosphate, a known intermediate of the de novo biosynthetic pathway. The existence of an adenosylcobinamide amidohydrolase enzyme would explain the lack of an adenosylcobinamide kinase in archaea.

Amidohydrolases↗