Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,675 records · Page 93Linked to original sources

Forms and functions of human SDR enzymes.

Short-chain dehydrogenases/reductases (SDR) are defined by distinct, common sequence motifs but constitute a functionally heterogenous superfamily of enzymes. At present, well over 1600 members from all forms of life are annotated in databases. Using the defined sequence motifs as queries, 37 distinct human members of the SDR family can be retrieved. The functional assignments of these forms fall minimally into three main groups, enzymes involved in intermediary metabolism, enzymes participating in lipid hormone and mediator metabolism, and open reading frames (ORFs) of yet undeciphered function. This overview, prepared just before completion of the human genome project, gives the different human SDR forms and relates them to human diseases.

Alcohol Dehydrogenase↗

A human protein atlas for normal and cancer tissues based on antibody proteomics.

Antibody-based proteomics provides a powerful approach for the functional study of the human proteome involving the systematic generation of protein-specific affinity reagents. We used this strategy to construct a comprehensive, antibody-based protein atlas for expression and localization profiles in 48 normal human tissues and 20 different cancers. Here we report a new publicly available database containing, in the first version, approximately 400,000 high resolution images corresponding to more than 700 antibodies toward human proteins. Each image has been annotated by a certified pathologist to provide a knowledge base for functional studies and to allow queries about protein profiles in normal and disease tissues. Our results suggest it should be possible to extend this analysis to the majority of all human proteins thus providing a valuable tool for medical and biological research.

Antibodies↗

Drosophila melanogaster as a model for studying protein-encoding genes that are resident in constitutive heterochromatin.

The organization of chromosomes into euchromatin and heterochromatin is one of the most enigmatic aspects of genome evolution. For a long time, heterochromatin was considered to be a genomic wasteland, incompatible with gene expression. However, recent studies--primarily conducted in Drosophila melanogaster--have shown that this peculiar genomic component performs important cellular functions and carries essential genes. New research on the molecular organization, function and evolution of heterochromatin has been facilitated by the sequencing and annotation of heterochromatic DNA. About 450 predicted genes have been identified in the heterochromatin of D. melanogaster, indicating that the number of active genes is higher than had been suggested by genetic analysis. Most of the essential genes are still unknown at the molecular level, and a detailed functional analysis of the predicted genes is difficult owing to the lack of mutant alleles. Far from being a peculiarity of Drosophila, heterochromatic genes have also been found in Saccharomyces cerevisiae, Schizosaccharomyces pombe, Oryza sativa and Arabidopsis thaliana, as well as in humans. The presence of expressed genes in heterochromatin seems paradoxical because they appear to function in an environment that has been considered incompatible with gene expression. In the future, genetic, functional genomic and proteomic analyses will offer powerful approaches with which to explore the functions of heterochromatic genes and to elucidate the mechanisms driving their expression.

Animals↗

The challenges and rewards of integrating diverse neuroscience information.

The design of database models and schemas for storing, cross-referencing, and retrieving neuroscience information faces issues that are similar but more complex than most of the other biomedical disciplines, such as genomics and proteonomics. Specifically, the visualization and manipulation of very large and diverse image data, such as digital brain atlases and functional magnetic resonance images, play a unique role in neuroscience while much of the associated information is textually recorded. Nongraphical information can include the annotation of large brain structures ranging from anatomical regions to intracellular structures, the description of cellular functional properties, and their various interrelationships, such as fiber connections. It is necessary that the heterogeneous and distributed types of data be cross-referenced to each other so that this diverse information can be efficiently retrieved, shared, and exchanged among the different neuroscientific disciplines. Continued advances in computers and Internet technologies appear to indicate that increasingly large data sets will be maintained on local or regional file servers and that informational interoperability will be achieved using a networked information system infrastructure. The authors and others have proposed and implemented models of semantically organized information systems that utilize centrally stored and highly structured archival information to index, cross-reference, and retrieve diverse, Web-based data sets.

Brain↗

Novel protein families in archaean genomes.

In a quest for novel functions in archaea, all archaean hypothetical open reading frames (ORFs), as annotated in the Swiss-Prot protein sequence database, were used to search the latest databases for the identification of characterized homologues. Of the 95 hypothetical archaean ORFs, 25 were found to be homologous to another hypothetical archaean ORF, while 36 were homologous to non-archaean proteins, of which as many as 30 were homologous to a characterized protein family. Thus the level of sequence similarity in this set reaches 64%, while the level of function assignment is only 32%. Of the ORFs with predicted functions, 12 homologies are reported here for the first time and represent nine new functions and one gene duplication at an acetyl-coA synthetase locus. The novel functions include components of the transcriptional and translational apparatus, such as ribosomal proteins, modification enzymes and a translation initiation factor. In addition, new enzymes are identified in archaea, such as cobyric acid synthase, dCTP deaminase and the first archaean homologues of a new subclass of ATP binding proteins found in fungi. Finally, it is shown that the putative laminin receptor family of eukaryotes and an archaean homologue belong to the previously characterized ribosomal protein family S2 from eubacteria. From the present and previous work, the major implication is that archaea seem to have a mode of expression of genetic information rather similar to eukaryotes, while eubacteria may have proceeded into unique ways of transcription and translation. In addition, with the detection of proteins in various metabolic and genetic processes in archaea, we can further predict the presence of additional proteins involved in these processes.

Animal Population Groups↗

Towards understanding the first genome sequence of a crenarchaeon by genome annotation using clusters of orthologous groups of proteins (COGs).

BACKGROUND: Standard archival sequence databases have not been designed as tools for genome annotation and are far from being optimal for this purpose. We used the database of Clusters of Orthologous Groups of proteins (COGs) to reannotate the genomes of two archaea, Aeropyrum pernix, the first member of the Crenarchaea to be sequenced, and Pyrococcus abyssi. RESULTS: A. pernix and P. abyssi proteins were assigned to COGs using the COGNITOR program; the results were verified on a case-by-case basis and augmented by additional database searches using the PSI-BLAST and TBLASTN programs. Functions were predicted for over 300 proteins from A. pernix, which could not be assigned a function using conventional methods with a conservative sequence similarity threshold, an approximately 50% increase compared to the original annotation. A. pernix shares most of the conserved core of proteins that were previously identified in the Euryarchaeota. Cluster analysis or distance matrix tree construction based on the co-occurrence of genomes in COGs showed that A. pernix forms a distinct group within the archaea, although grouping with the two species of Pyrococci, indicative of similar repertoires of conserved genes, was observed. No indication of a specific relationship between Crenarchaeota and eukaryotes was obtained in these analyses. Several proteins that are conserved in Euryarchaeota and most bacteria are unexpectedly missing in A. pernix, including the entire set of de novo purine biosynthesis enzymes, the GTPase FtsZ (a key component of the bacterial and euryarchaeal cell-division machinery), and the tRNA-specific pseudouridine synthase, previously considered universal. A. pernix is represented in 48 COGs that do not contain any euryarchaeal members. Many of these proteins are TCA cycle and electron transport chain enzymes, reflecting the aerobic lifestyle of A. pernix. CONCLUSIONS: Special-purpose databases organized on the basis of phylogenetic analysis and carefully curated with respect to known and predicted protein functions provide for a significant improvement in genome annotation. A differential genome display approach helps in a systematic investigation of common and distinct features of gene repertoires and in some cases reveals unexpected connections that may be indicative of functional similarities between phylogenetically distant organisms and of lateral gene exchange.

Archaea↗

Association of genes to genetically inherited diseases using data mining.

Although approximately one-quarter of the roughly 4,000 genetically inherited diseases currently recorded in respective databases (LocusLink, OMIM) are already linked to a region of the human genome, about 450 have no known associated gene. Finding disease-related genes requires laborious examination of hundreds of possible candidate genes (sometimes, these are not even annotated; see, for example, refs 3,4). The public availability of the human genome draft sequence has fostered new strategies to map molecular functional features of gene products to complex phenotypic descriptions, such as those of genetically inherited diseases. Owing to recent progress in the systematic annotation of genes using controlled vocabularies, we have developed a scoring system for the possible functional relationships of human genes to 455 genetically inherited diseases that have been mapped to chromosomal regions without assignment of a particular gene. In a benchmark of the system with 100 known disease-associated genes, the disease-associated gene was among the 8 best-scoring genes with a 25% chance, and among the best 30 genes with a 50% chance, showing that there is a relationship between the score of a gene and its likelihood of being associated with a particular disease. The scoring also indicates that for some diseases, the chance of identifying the underlying gene is higher.

Chromosome Mapping↗

A molecular profile of the mouse gastric parietal cell with and without exposure to Helicobacter pylori.

The parietal cell (PC) plays an important role in normal gastric physiology and in common diseases of the stomach. Although the genes involved in acid secretion are well known, there is limited molecular information about other aspects of PC function. We have generated a comprehensive database of genes expressed preferentially in PCs relative to other gastric mucosal cell lineages. PCs were purified from FVB/N mouse stomachs by lectin panning. cRNA generated from PC-enriched (PC(+)) and PC-depleted (PC(-)) populations were used to query oligonucleotide-based microarrays. False-positive signals were filtered by using a new algorithm for noise reduction and selected results independently audited by real-time quantitative reverse transcription (RT)-PCR. The annotated database of 240 genes reveals previously unappreciated aspects of cellular function, including factors that may mediate PC regulation of gastric stem cell proliferation. PC(+) and PC(-) expression profiles were also prepared from germ-free mice 2 and 8 weeks after colonization with a clinical isolate of Helicobacter pylori (Hp)--the pathogen that produces acid-peptic disease (gastritis, ulcers) in humans. Whereas PC(+) gene expression was remarkably constant, the PC(-) fractions demonstrated a robust, evolving host response, with increased expression of genes involved in cell motility/migration, extracellular matrix interactions, and IFN responses. The consistency of PC(+) gene expression allowed identification of a cohort of 92 genes enriched in PCs under all conditions studied. These genes provide a molecular profile that can be used to define this epithelial lineage under a variety of physiologic, pharmacologic, and pathologic stimuli.

Animals↗

Detection of weakly conserved ancestral mammalian regulatory sequences by primate comparisons.

BACKGROUND: Genomic comparisons between human and distant, non-primate mammals are commonly used to identify cis-regulatory elements based on constrained sequence evolution. However, these methods fail to detect functional elements that are too weakly conserved among mammals to distinguish them from non-functional DNA. RESULTS: To evaluate a strategy for large scale genome annotation that is complementary to the commonly used distal species comparisons, we explored the potential of deep intra-primate sequence comparisons. We sequenced the orthologs of 558 kb of human genomic sequence, covering multiple loci involved in cholesterol homeostasis, in 6 non-human primates. Our analysis identified six non-coding DNA elements displaying significant conservation among primates but undetectable in more distant comparisons. In vitro and in vivo tests revealed that at least three of these six elements have regulatory function. Notably, the mouse orthologs of these three functional human sequences had regulatory activity despite their lack of significant sequence conservation, indicating that they are ancestral mammalian cis-regulatory elements. These regulatory elements could be detected even in a smaller set of three primate species including human, rhesus and marmoset. CONCLUSION: We have demonstrated that intra-primate sequence comparisons can be used to identify functional modules in large genomic regions, including cis-regulatory elements that are not detectable through comparison with non-mammalian genomes. With the available human and rhesus genomes and that of marmoset, which is being actively sequenced, this strategy can be extended to the whole genome in the near future.

Animals↗

Sequence-based heuristics for faster annotation of non-coding RNA families.

MOTIVATION: Non-coding RNAs (ncRNAs) are functional RNA molecules that do not code for proteins. Covariance Models (CMs) are a useful statistical tool to find new members of an ncRNA gene family in a large genome database, using both sequence and, importantly, RNA secondary structure information. Unfortunately, CM searches are extremely slow. Previously, we created rigorous filters, which provably sacrifice none of a CM's accuracy, while making searches significantly faster for virtually all ncRNA families. However, these rigorous filters make searches slower than heuristics could be. RESULTS: In this paper we introduce profile HMM-based heuristic filters. We show that their accuracy is usually superior to heuristics based on BLAST. Moreover, we compared our heuristics with those used in tRNAscan-SE, whose heuristics incorporate a significant amount of work specific to tRNAs, where our heuristics are generic to any ncRNA. Performance was roughly comparable, so we expect that our heuristics provide a high-quality solution that--unlike family-specific solutions--can scale to hundreds of ncRNA families. AVAILABILITY: The source code is available under GNU Public License at the supplementary web site.

Algorithms↗

Clustering biological annotations and gene expression data to identify putatively co-regulated biological processes.

MOTIVATION: Functional profiling is a key step of microarray gene expression data analysis. Identifying co-regulated biological processes could help for better understanding of underlying biological interactions within the studied biological frame. RESULTS: We present herein an original approach designed to search for putatively co-regulated biological processes sharing a significant number of co-expressed genes. An R language implementation named "FunCluster" was built and tested on two gene expression data sets. A discriminatory functional analysis of the first data set, related to experiments performed on separated adipocytes and stroma vascular fraction cells of human white adipose tissue, highlighted the prevalent role of nonadipose cells in the synthesis of inflammatory and immunity molecules in human adiposity. On the second data set, resulting from a model investigating insulin coordinated regulation of gene expression in human skeletal muscle, FunCluster analysis spotlighted novel functional classes of putatively co-regulated biological processes related to protein metabolism and the regulation of muscular contraction. AVAILABILITY: Supplementary information about the FunCluster tool is available on-line at http://corneliu.henegar.info/FunCluster.htm.

Algorithms↗

Mining sequence annotation databanks for association patterns.

MOTIVATION: Millions of protein sequences currently being deposited to sequence databanks will never be annotated manually. Similarity-based annotation generated by automatic software pipelines unavoidably contains spurious assignments due to the imperfection of bioinformatics methods. Examples of such annotation errors include over- and underpredictions caused by the use of fixed recognition thresholds and incorrect annotations caused by transitivity based information transfer to unrelated proteins or transfer of errors already accumulated in databases. One of the most difficult and timely challenges in bioinformatics is the development of intelligent systems aimed at improving the quality of automatically generated annotation. A possible approach to this problem is to detect anomalies in annotation items based on association rule mining. RESULTS: We present the first large-scale analysis of association rules derived from two large protein annotation databases-Swiss-Prot and PEDANT-and reveal novel, previously unknown tendencies of rule strength distributions. Most of the rules are either very strong or very weak, with rules in the medium strength range being relatively infrequent. Based on dynamics of error correction in subsequent Swiss-Prot releases and on our own manual analysis we demonstrate that exceptions from strong rules are, indeed, significantly enriched in annotation errors and can be used to automatically flag them. We identify different strength dependencies of rules derived from different fields in Swiss-Prot. A compositional breakdown of association rules generated from PEDANT in terms of their constituent items indicates that most of the errors that can be corrected are related to gene functional roles. Swiss-Prot errors are usually caused by under-annotation owing to its conservative approach, whereas automatically generated PEDANT annotation suffers from over-annotation. AVAILABILITY: All data generated in this study are available for download and browsing at http://pedant.gsf.de/ARIA/index.htm.

Conserved Sequence↗

GeomeTRe: accurate calculation of geometrical descriptors of tandem repeat proteins.

MOTIVATION: Structured tandem repeat proteins (STRPs) are characterized by preserved structural motifs arranged in a modular way. The structural and functional diversity of STRPs makes them particularly important for studying evolution and novel structure-function relationships, and ultimately for designing new synthetic proteins with specific functions. One crucial aspect of their classification is the estimation of geometrical parameters, which can provide better insight into their properties and the relationship between the spatial arrangement of repeated units and protein function. Calculating geometric descriptors for STRPs is challenging because naturally occurring repeats are not "perfect" and often contain insertions and deletions. Existing tools for predicting structural symmetry work well on simple cases but often fail for most natural proteins. RESULTS: Here, we present GeomeTRe, an algorithm that calculates geometrical descriptors such as curvature (yaw), twist (roll), and pitch for a protein structure with known repeat unit positions. The algorithm simulates the movement of consecutive units, identifies rotational axes, and calculates the corresponding Tait-Bryan angles. GeomeTRe's parameters can enhance STRP annotation and classification by identifying variations in geometric arrangements among different functional groups. The package is fast and suitable for processing large protein structure datasets when repeat region information (e.g. from RepeatsDB) is available. AVAILABILITY AND IMPLEMENTATION: GeomeTRe is available as a Python package; source code and documentation can be found at https://github.com/BioComputingUP/GeomeTRe.

Algorithms↗

Mapping of conserved RNA secondary structures predicts thousands of functional noncoding RNAs in the human genome.

In contrast to the fairly reliable and complete annotation of the protein coding genes in the human genome, comparable information is lacking for noncoding RNAs (ncRNAs). We present a comparative screen of vertebrate genomes for structural noncoding RNAs, which evaluates conserved genomic DNA sequences for signatures of structural conservation of base-pairing patterns and exceptional thermodynamic stability. We predict more than 30,000 structured RNA elements in the human genome, almost 1,000 of which are conserved across all vertebrates. Roughly a third are found in introns of known genes, a sixth are potential regulatory elements in untranslated regions of protein-coding mRNAs and about half are located far away from any known gene. Only a small fraction of these sequences has been described previously. A comparison with recent tiling array data shows that more than 40% of the predicted structured RNAs overlap with experimentally detected sites of transcription. The widespread conservation of secondary structure points to a large number of functional ncRNAs and cis-acting mRNA structures in the human genome.

Animals↗

The mouse genome database (MGD): new features facilitating a model system.

The mouse genome database (MGD, http://www.informatics.jax.org/), the international community database for mouse, provides access to extensive integrated data on the genetics, genomics and biology of the laboratory mouse. The mouse is an excellent and unique animal surrogate for studying normal development and disease processes in humans. Thus, MGD's primary goals are to facilitate the use of mouse models for studying human disease and enable the development of translational research hypotheses based on comparative genotype, phenotype and functional analyses. Core MGD data content includes gene characterization and functions, phenotype and disease model descriptions, DNA and protein sequence data, polymorphisms, gene mapping data and genome coordinates, and comparative gene data focused on mammals. Data are integrated from diverse sources, ranging from major resource centers to individual investigator laboratories and the scientific literature, using a combination of automated processes and expert human curation. MGD collaborates with the bioinformatics community on the development of data and semantic standards, and it incorporates key ontologies into the MGD annotation system, including the Gene Ontology (GO), the Mammalian Phenotype Ontology, and the Anatomical Dictionary for Mouse Development and the Adult Anatomy. MGD is the authoritative source for mouse nomenclature for genes, alleles, and mouse strains, and for GO annotations to mouse genes. MGD provides a unique platform for data mining and hypothesis generation where one can express complex queries simultaneously addressing phenotypic effects, biochemical function and process, sub-cellular location, expression, sequence, polymorphism and mapping data. Both web-based querying and computational access to data are provided. Recent improvements in MGD described here include the incorporation of single nucleotide polymorphism data and search tools, the addition of PIR gene superfamily classifications, phenotype data for NIH-acquired knockout mice, images for mouse phenotypic genotypes, new functional graph displays of GO annotations, and new orthology displays including sequence information and graphic displays.

Animals↗

Identification of function-associated loop motifs and application to protein function prediction.

MOTIVATION: The detection of function-related local 3D-motifs in protein structures can provide insights towards protein function in absence of sequence or fold similarity. Protein loops are known to play important roles in protein function and several loop classifications have been described, but the automated identification of putative functional 3D-motifs in such classifications has not yet been addressed. This identification can be used on sequence annotations. RESULTS: We evaluated three different scoring methods for their ability to identify known motifs from the PROSITE database in ArchDB. More than 500 new putative function-related motifs not reported in PROSITE were identified. Sequence patterns derived from these motifs were especially useful at predicting precise annotations. The number of reliable sequence annotations could be increased up to 100% with respect to standard BLAST. CONTACT: boliva@imim.es SUPPLEMENTARY INFORMATION: Supplementary Data are available at Bioinformatics online.

Amino Acid Sequence↗

Fold independent structural comparisons of protein-ligand binding sites for exploring functional relationships.

The rapid growth in protein structural data and the emergence of structural genomics projects have increased the need for automatic structure analysis and tools for function prediction. Small molecule recognition is critical to the function of many proteins; therefore, determination of ligand binding site similarity is important for understanding ligand interactions and may allow their functional classification. Here, we present a binding sites database (SitesBase) that given a known protein-ligand binding site allows rapid retrieval of other binding sites with similar structure independent of overall sequence or fold similarity. However, each match is also annotated with sequence similarity and fold information to aid interpretation of structure and functional similarity. Similarity in ligand binding sites can indicate common binding modes and recognition of similar molecules, allowing potential inference of function for an uncharacterised protein or providing additional evidence of common function where sequence or fold similarity is already known. Alternatively, the resource can provide valuable information for detailed studies of molecular recognition including structure-based ligand design and in understanding ligand cross-reactivity. Here, we show examples of atomic similarity between superfamily or more distant fold relatives as well as between seemingly unrelated proteins. Assignment of unclassified proteins to structural superfamiles is also undertaken and in most cases substantiates assignments made using sequence similarity. Correct assignment is also possible where sequence similarity fails to find significant matches, illustrating the potential use of binding site comparisons for newly determined proteins.

Animals↗

The Role of Small Segmental Duplications in Generating Identical Isoforms Through Alternative Splicing Sites.

Alternative splicing plays a crucial role in expanding proteomic diversity but can also generate identical isoforms under certain conditions. While mutually exclusive splicing of tandem exons has occasionally been reported to produce identical isoforms, the extent to which other splicing events contribute to this phenomenon remains unclear. In this study, we demonstrate that alternative 5' and 3' splice site selection can also lead to the formation of identical isoforms, providing an additional type of splicing event for functional redundancy in transcriptomes. To address this, we analyzed reference genome annotations from 15 plant species, including Arabidopsis thaliana and wheat (Triticum aestivum), obtained from the RefSeq database. Identical isoforms were computationally defined as transcripts with distinct exon-intron structures but identical coding sequences. Our analysis reveals that the majority of alternative 5' and 3' fragments originate from small segmental duplications, suggesting that sequence repetition within gene regions facilitates the emergence of such splicing patterns. We also observed differences in the annotated 5' UTRs of some identical isoforms. However, since the alternative splicing sites themselves were not located within UTRs, these differences may reflect annotation uncertainty rather than genuine AS-derived variation. Given that UTR predictions in reference databases are not always precise, such observations should be interpreted cautiously. Expression analysis using an isoform-specific k-mer approach confirmed that identical isoforms can be differentially regulated. These findings suggest that, beyond expanding protein diversity, alternative splicing can also generate redundant isoforms that are differentially expressed at the RNA level, indicating potential regulatory roles. By elucidating the structural and regulatory factors contributing to the formation and retention of identical isoforms, our study provides new insights into the evolutionary and functional significance of alternative splicing in plants.

Alternative Splicing↗