Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,639 records · Page 91Linked to original sources

SMART: identification and annotation of domains from signalling and extracellular protein sequences.

SMART is a simple modular architecture research tool and database that provides domain identification and annotation on the WWW (http://coot.embl-heidelberg.de/SMART). The tool compares query sequences with its databases of domain sequences and multiple alignments whilst concurrently identifying compositionally biased regions such as signal peptide, transmembrane and coiled coil segments. Annotated and unannotated regions of the sequence can be used as queries in searches of sequence databases. The SMART alignment collection represents more than 250 signalling and extracellular domains. Each alignment is curated to assign appropriate domain boundaries and to ensure its quality. In addition, each domain is annotated extensively with respect to cellular localisation, species distribution, functional class, tertiary structure and functionally important residues.

Amino Acid Sequence↗

An optimized protocol for analysis of EST sequences.

The vast body of Expressed Sequence Tag (EST) data in the public databases provide an important resource for comparative and functional genomics studies and an invaluable tool for the annotation of genomic sequences. We have developed a rigorous protocol for reconstructing the sequences of transcribed genes from EST and gene sequence fragments. A key element in developing this protocol has been the evaluation of a number of sequence assembly programs to determine which most faithfully reproduce transcript sequences from EST data. The TIGR Gene Indices constructed using this protocol for human, mouse, rat and a variety of other plant and animal models have demonstrated their utility in a variety of applications and are freely available to the scientific research community.

Algorithms↗

PatSearch: A program for the detection of patterns and structural motifs in nucleotide sequences.

Regulation of gene expression at transcriptional and post-transcriptional level involves the interaction between short DNA or RNA tracts and the corresponding trans-acting protein factors. Detection of such cis-acting elements in genome-wide screenings may significantly contribute to genome annotation and comparative analysis as well as to target functional characterization experiments. We present here PatSearch, a flexible and fast pattern matcher able to search for specific combinations of oligonucleotide consensus sequences, secondary structure elements and position-weight matrices. It can also allow for mismatches/mispairings below a user fixed threshold. We report three different applications of the program in the search of complex patterns such as those of the iron responsive element hairpin-loop structure, the p53 responsive element and a promoter module containing CAAT-, TATA- and cap-boxes. PatSearch is available on the web at http://bighost.area.ba.cnr.it/BIG/PatSearch/.

Base Sequence↗

The Genexpress Index: a resource for gene discovery and the genic map of the human genome.

Detailed analysis of a set of 18,698 sequences derived from both ends of 10,979 human skeletal muscle and brain cDNA clones defined 6676 functional families, characterized by their sequence signatures over 5750 distinct human gene transcripts. About half of these genes have been assigned to specific chromosomes utilizing 2733 eSTS markers, the polymerase chain reaction, and DNA from human-rodent somatic cell hybrids. Sequence and clone clustering and a functional classification together with comprehensive data base searches and annotations made it possible to develop extensive sequence and map cross-indexes, define electronic expression profiles, identify a new set of overlapping genes, and provide numerous new candidate genes for human pathologies.

Amino Acid Sequence↗

MatrixExplorer: a dual-representation system to explore social networks.

MatrixExplorer is a network visualization system that uses two representations: node-link diagrams and matrices. Its design comes from a list of requirements formalized after several interviews and a participatory design session conducted with social science researchers. Although matrices are commonly used in social networks analysis, very few systems support the matrix-based representations to visualize and analyze networks. MatrixExplorer provides several novel features to support the exploration of social networks with a matrix-based representation, in addition to the standard interactive filtering and clustering functions. It provides tools to reorder (layout) matrices, to annotate and compare findings across different layouts and find consensus among several clusterings. MatrixExplorer also supports Node-link diagram views which are familiar to most users and remain a convenient way to publish or communicate exploration results. Matrix and node-link representations are kept synchronized at all stages of the exploration process.

Algorithms↗

Neuroendocrine-immune interactions in homeostasis and autoimmunity.

Recent experimental evidence confirms the interrelationships between the central nervous, neuroendocrine and immune systems. Indeed, extensive duality exists in the use of neurotransmitters, hormones and receptors each system displays. In the present annotation, the effect of cytokines, soluble mediators of immune function, on the CNS and neuroendocrine systems is addressed and conversely, we discuss the modification of the immune compartment by the sympathetic nervous and neuroendocrine systems, with particular reference to the role of noradrenaline and corticosterone. Dysfunction between the systems is considered in the context of autoimmune conditions, with emphasis on experimental allergic encephalomyelitis and the contribution of corticosterone-driven T-cell apoptosis to recovery from the disease. Finally, we speculate on the relevance of neuroimmune interactions in the pathogenesis of multiple sclerosis.

Animals↗

Systematic identification of selective essential genes in Helicobacter pylori by genome prioritization and allelic replacement mutagenesis.

A comparative genomic approach was used to identify Helicobacter pylori 26695 open reading frames (ORFs) which are conserved in H. pylori J99 but highly diverged in other eubacteria. A survey of selected pathways of central intermediary metabolism was also carried out, and genes with a potentially selective role in H. pylori were identified. Forty-five ORFs identified in these two analyses were screened using a rapid vector-free allelic replacement mutagenesis technique, and 33 were shown to be essential in vitro. Notably, 13 ORFs gave essentiality results which are unexpected in view of their known or proposed functions, and phylogenetic analysis was used to investigate the annotation of 7 such ORFs which are highly diverged. We propose that the products of a number of these H. pylori-specific essential genes may be suitable targets for novel anti-H. pylori therapies.

Alleles↗

Molecular models of NS3 protease variants of the Hepatitis C virus.

BACKGROUND: Hepatitis C virus (HCV) currently infects approximately three percent of the world population. In view of the lack of vaccines against HCV, there is an urgent need for an efficient treatment of the disease by an effective antiviral drug. Rational drug design has not been the primary way for discovering major therapeutics. Nevertheless, there are reports of success in the development of inhibitor using a structure-based approach. One of the possible targets for drug development against HCV is the NS3 protease variants. Based on the three-dimensional structure of these variants we expect to identify new NS3 protease inhibitors. In order to speed up the modeling process all NS3 protease variant models were generated in a Beowulf cluster. The potential of the structural bioinformatics for development of new antiviral drugs is discussed. RESULTS: The atomic coordinates of crystallographic structure 1CU1 and 1DY9 were used as starting model for modeling of the NS3 protease variant structures. The NS3 protease variant structures are composed of six subdomains, which occur in sequence along the polypeptide chain. The protease domain exhibits the dual beta-barrel fold that is common among members of the chymotrypsin serine protease family. The helicase domain contains two structurally related beta-alpha-beta subdomains and a third subdomain of seven helices and three short beta strands. The latter domain is usually referred to as the helicase alpha-helical subdomain. The rmsd value of bond lengths and bond angles, the average G-factor and Verify 3D values are presented for NS3 protease variant structures. CONCLUSIONS: This project increases the certainty that homology modeling is an useful tool in structural biology and that it can be very valuable in annotating genome sequence information and contributing to structural and functional genomics from virus. The structural models will be used to guide future efforts in the structure-based drug design of a new generation of NS3 protease variants inhibitors. All models in the database are publicly accessible via our interactive website, providing us with large amount of structural models for use in protein-ligand docking analysis.

Amino Acid Sequence↗

Profiling of pathway-specific changes in gene expression following growth of human cancer cell lines transplanted into mice.

BACKGROUND: Tumor cells cultured in vitro are widely used to investigate the molecular biology of cancers and to evaluate responses to drugs and other agents. The full extent to which gene expression in cancer cells is modulated by extrinsic factors and by the microenvironment in which the cancer cells reside remains to be determined. Two cancer cell lines (A549 lung adenocarcinoma and U118 glioblastoma) were transplanted subcutaneously into immunodeficient mice to form tumors. Global gene-expression profiles of the tumors were determined, based on analysis of expression of human genes, and compared with expression profiles of the cell lines grown in culture. RESULTS: A bioinformatics approach associated genes that showed changes in their expression levels with functional classes as defined by either the GO gene annotations or MeSH terms in the literature. The classes of genes expressed at higher levels in cells grown in vitro indicated increased cell division and metabolism, reflecting the more favorable environment for cell proliferation. In contrast, in vivo tumor growth resulted in upregulation of a significant number of genes involved in the extracellular matrix (ECM), cell adhesion, cytokine and metalloendopeptidase activity, and neovascularization. When placed in comparable tissue environments, the U118 cells and the A549 cells expressed different sets of ECM and cell adhesion-related genes, suggesting different mechanisms of extracellular interaction at work in the different cancers. CONCLUSIONS: Studies of this type allow us to examine the specific contribution of cancer cells to gene expression patterns within an in vivo tumor mixed with non-cancerous tissue.

Animals↗

Phylogenetic and functional classification of ATP-binding cassette (ABC) systems.

ATP binding cassette (ABC) systems constitute one of the most abundant superfamilies of proteins. They are involved in the transport of a wide variety of substances, but also in many cellular processes and in their regulation. In this paper, we made a comparative analysis of the properties of ABC systems and we provide a phylogenetic and functional classification. This analysis will be helpful to accurately annotate ABC systems discovered during the sequencing of the genome of living organisms and to identify the partners of the ABC ATPases.

ATP-Binding Cassette Transporters↗

The SBASE protein domain library, release 3.0: a collection of annotated protein sequence segments.

SBASE 3.0 is the third release of SBASE, a collection of annotated protein domain sequences. SBASE entries represent various structural, functional, ligand-binding and topogenic segments of proteins as defined by their publishing authors. SBASE can be used for establishing domain homologies using different database-search tools such as FASTA [Lipman and Pearson (1985) Science, 227, 1436-1441], and BLAST3 [Altschul and Lipman (1990) Proc. Natl. Acad. Sci. USA, 87, 5509-5513] which is especially useful in the case of loosely defined domain types for which efficient consensus patterns can not be established. The present release contains 41,749 entries provided with standardized names and cross-referenced to the major protein and nucleic acid databanks as well as to the PROSITE catalogue of protein sequence patterns. The entries are clustered into 2285 groups using the BLAST algorithm for computing similarity measures. SBASE 3.0 is freely available on request to the authors or by anonymous 'ftp' file transfer from < ftp.icgeb.trieste.it >. Individual records can be retrieved with the gopher server at < icgeb.trieste.it > and with a www-server at < http:@www.icgeb.trieste.it >. Automated searching of SBASE by BLAST can be carried out with the electronic mail server < sbase@icgeb.trieste.it >. Another mail server < domain@hubi.abc.hu > assigns SBASE domain homologies on the basis of SWISS-PROT searches. A comparison of pertinent search strategies is presented.

Amino Acid Sequence↗

InterPro--an integrated documentation resource for protein families, domains and functional sites.

MOTIVATION: InterPro is a new integrated documentation resource for protein families, domains and functional sites, developed initially as a means of rationalising the complementary efforts of the PROSITE, PRINTS, Pfam and ProDom database projects. RESULTS: Merged annotations from PRINTS, PROSITE and Pfam form the InterPro core. Each combined InterPro entry includes functional descriptions and literature references, and links are made back to the relevant parent database(s), allowing users to see at a glance whether a particular family or domain has associated patterns, profiles, fingerprints, etc. Merged and individual entries (i.e. those that have no counterpart in the companion resources) are assigned unique accession numbers. Release 1.2 of InterPro (June 2000) contains over 3000 entries, representing families, domains, repeats and sites of post-translational modification (PTMs) encoded by 6581 different regular expressions, profiles, fingerprints and Hidden Markov Models (HMMs). Each InterPro entry lists all the matches against SWISS-PROT and TrEMBL (more than 1000000 hits from 264333 different proteins out of 384572 in SWISS-PROT and TrEMBL).

Computational Biology↗

PCAS--a precomputed proteome annotation database resource.

BACKGROUND: Many model proteomes or "complete" sets of proteins of given organisms are now publicly available. Much effort has been invested in computational annotation of those "draft" proteomes. Motif or domain based algorithms play a pivotal role in functional classification of proteins. Employing most available computational algorithms, mainly motif or domain recognition algorithms, we set up to develop an online proteome annotation system with integrated proteome annotation data to complement existing resources. RESULTS: We report here the development of PCAS (ProteinCentric Annotation System) as an online resource of pre-computed proteome annotation data. We applied most available motif or domain databases and their analysis methods, including hmmpfam search of HMMs in Pfam, SMART and TIGRFAM, RPS-PSIBLAST search of PSSMs in CDD, pfscan of PROSITE patterns and profiles, as well as PSI-BLAST search of SUPERFAMILY PSSMs. In addition, signal peptide and TM are predicted using SignalP and TMHMM respectively. We mapped SUPERFAMILY and COGs to InterPro, so the motif or domain databases are integrated through InterPro. PCAS displays table summaries of pre-computed data and a graphical presentation of motifs or domains relative to the protein. As of now, PCAS contains human IPI, mouse IPI, and rat IPI, A. thaliana, C. elegans, D. melanogaster, S. cerevisiae, and S. pombe proteome.PCAS is available at http://pak.cbi.pku.edu.cn/proteome/gca.php CONCLUSION: PCAS gives better annotation coverage for model proteomes by employing a wider collection of available algorithms. Besides presenting the most confident annotation data, PCAS also allows customized query so users can inspect statistically less significant boundary information as well. Therefore, besides providing general annotation information, PCAS could be used as a discovery platform. We plan to update PCAS twice a year. We will upgrade PCAS when new proteome annotation algorithms identified.

Algorithms↗

An expression system for the functional analysis of pheromone genes in the tetrapolar basidiomycete Schizophyllum commune.

The investigation of putative pheromone genes of basidiomycetes has been difficult since the small open reading frames are essentially annotated on the basis of a C-terminal farnesylation signal. In order to identify the functional reading frame, expression of small DNA fragments in the fungus is necessary. The expression system developed in the presented paper allows fusion to the promoter of the tef1 gene encoding the constitutively and highly expressed translation elongation factor EF1alpha. This system has been shown to be functional using an easily selectable gene, ura1. The application to identification of functional pheromone genes has been shown with the newly detected bap2(4) gene. The Bap2(4) pheromone is the first Balpha pheromone gene activating only a single receptor specificity.

Amino Acid Sequence↗

Computational analyses and annotations of the Arabidopsis peroxidase gene family.

Classical heme-containing plant peroxidases have been ascribed a wide variety of functional roles related to development, defense, lignification, and hormonal signaling. More than 40 peroxidase genes are now known in Arabidopsis thaliana for which functional association is complicated by a general lack of peroxidase substrate specificity. Computational analysis was performed on 30 near full-length Arabidopsis peroxidase cDNAs for annotation of start codons and signal peptide cleavage sites. A compositional analysis revealed that 23 of the 30 peroxidase cDNAs have 5' untranslated regions containing 40-71% adenine, a rare feature observed also in cDNAs which predominantly encode stress-induced proteins, and which may indicate translational regulation.

Adenine↗

Multi-species sequence comparison: the next frontier in genome annotation.

Multi-species comparisons of DNA sequences are more powerful for discovering functional sequences than pairwise DNA sequence comparisons. Most current computational tools have been designed for pairwise comparisons, and efficient extension of these tools to multiple species will require knowledge of the ideal evolutionary distance to choose and the development of new algorithms for alignment, analysis of conservation, and visualization of results.

Animals↗

A genomics approach to crop pest and disease research.

Genome-wide analyses of gene function and gene expression are beginning to yield valuable information in many areas of biological research, and these genomic tools are now being applied to crop pest and disease research. DNA sequencing of cDNA libraries to generate sets of expressed sequence tags (ESTs) are allowing gene compendiums for crop diseases to be compiled. Annotation of such data collections is also providing a wealth of functional information about gene products through similarities to proteins with known function. The next phase of the functional genomics era will be to employ large-scale techniques to knock out or silence genes in order to synthesize gene-specific mutants for phenotypic analysis and to use micro-array methodology to analyze global gene expression, protein turnover and protein processing during the processes of parasitism and colonization. Application of these technologies promises to accelerate the pace that biological information relevant to crop protection accrues. The ability of researchers to assimilate this information into complex models and workable hypotheses is, thus, set to revolutionize the way we study pests and diseases of crop plants.

Crops, Agricultural↗

MPact: the MIPS protein interaction resource on yeast.

In recent years, the Munich Information Center for Protein Sequences (MIPS) yeast protein-protein interaction (PPI) dataset has been used in numerous analyses of protein networks and has been called a gold standard because of its quality and comprehensiveness [H. Yu, N. M. Luscombe, H. X. Lu, X. Zhu, Y. Xia, J. D. Han, N. Bertin, S. Chung, M. Vidal and M. Gerstein (2004) Genome Res., 14, 1107-1118]. MPact and the yeast protein localization catalog provide information related to the proximity of proteins in yeast. Beside the integration of high-throughput data, information about experimental evidence for PPIs in the literature was compiled by experts adding up to 4300 distinct PPIs connecting 1500 proteins in yeast. As the interaction data is a complementary part of CYGD, interactive mapping of data on other integrated data types such as the functional classification catalog [A. Ruepp, A. Zollner, D. Maier, K. Albermann, J. Hani, M. Mokrejs, I. Tetko, U. Güldener, G. Mannhaupt, M. Münsterkötter and H. W. Mewes (2004) Nucleic Acids Res., 32, 5539-5545] is possible. A survey of signaling proteins and comparison with pathway data from KEGG demonstrates that based on these manually annotated data only an extensive overview of the complexity of this functional network can be obtained in yeast. The implementation of a web-based PPI-analysis tool allows analysis and visualization of protein interaction networks and facilitates integration of our curated data with high-throughput datasets. The complete dataset as well as user-defined sub-networks can be retrieved easily in the standardized PSI-MI format. The resource can be accessed through http://mips.gsf.de/genre/proj/mpact.

Databases, Protein↗