Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Functional annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Analysis of the mouse transcriptome based on functional annotation of 60,770 full-length cDNAs.

Only a small proportion of the mouse genome is transcribed into mature messenger RNA transcripts. There is an international collaborative effort to identify all full-length mRNA transcripts from the mouse, and to ensure that each is represented in a physical collection of clones. Here we report the manual annotation of 60,770 full-length mouse complementary DNA sequences. These are clustered into 33,409 'transcriptional units', contributing 90.1% of a newly established mouse transcriptome database. Of these transcriptional units, 4,258 are new protein-coding and 11,665 are new non-coding messages, indicating that non-coding RNA is a major component of the transcriptome. 41% of all transcriptional units showed evidence of alternative splicing. In protein-coding transcripts, 79% of splice variations altered the protein product. Whole-transcriptome analyses resulted in the identification of 2,431 sense-antisense pairs. The present work, completely supported by physical clones, provides the most comprehensive survey of a mammalian transcriptome so far, and is a valuable resource for functional genomics.

Alternative Splicing↗

A novel genetic island of meningitic Escherichia coli K1 containing the ibeA invasion gene (GimA): functional annotation and carbon-source-regulated invasion of human brain microvascular endothelial cells.

The IbeA (ibe10) gene is an invasion determinant contributing to E. coli K1 invasion of the blood-brain barrier. This gene has been cloned and characterized from the chromosome of an invasive cerebrospinal fluid isolate of E. coli K1, strain RS218 (018:K1: H7). In the present study, a genetic island of meningitic E. coli containing ibeA (GimA) has been identified. A 20.3-kb genomic DNA island unique to E. coli K1 strains has been cloned and sequenced from an RS218 E. coli K1 genomic DNA library. Fourteen new genes have been identified in addition to the ibeA. The DNA sequence analysis indicated that the ibeA gene cluster was localized to the 98 min region and consisted of four operons, ptnIPKC, cglDTEC, gcxKRCI and ibeRAT. The G+C content (46.2%) of unique regions of the island is substantially different from that (50.8%) of the rest of the E. coli chromosome. By computer-assisted analysis of the sequences with DNA and protein databases (GenBank and PROSITE databases), the functions of the gene products could be anticipated, and were assigned to the functional categories of proteins relating to carbon source metabolism and substrate transportation. Glucose was shown to enhance E. coli penetration of human brain microvascular endothelial cells and exogenous cAMP was able to block the stimulating effect of glucose, suggesting that catabolic regulation may play a role in control of E. coli K1 invasion gene expression. Our data suggest that this genetic island may contribute to E. coli invasion of the blood-brain barrier through a carbon-source-regulated process.

Amino Acid Sequence↗

Functional annotation of a novel NFKB1 promoter polymorphism that increases risk for ulcerative colitis.

Nuclear Factor-kappaB (NF-kappaB) is a major transcription regulator of immune response, apoptosis and cell-growth control genes, and is upregulated in inflammatory bowel disease (IBD), both ulcerative colitis (UC) and Crohn's disease. The NFKB1 gene encodes the NF-kappaB p105/p50 isoforms. Genome-wide screens in IBD families show evidence for linkage on chromosome 4q where NFKB1 maps. We sequenced the NFKB1 promoter, exon 1 and all coding exons in 10 IBD probands and two controls, and identified six nucleotide variants, including a common insertion/deletion promoter polymorphism (-94ins/delATTG). Using pedigree-based transmission disequilibrium tests, we observed modest evidence for linkage disequilibrium (LD), independent of linkage, between the -94delATTG allele and UC in 131 out of 235 IBD pedigrees with UC offspring (P=0.047-0.052). This allele was also more frequent in the 156 non-Jewish UC probands from the 235 IBD pedigrees than in 149 non-Jewish controls (P=0.015). The -94delATTG association with UC was replicated in a second set of 258 unrelated, non-Jewish UC cases and 653 new, non-Jewish controls (P=0.021). Nuclear proteins from normal human colon tissue and colonic cell lines, but not ileal tissue, showed significant binding to -94insATTG but not to -94delATTG containing oligonucleotides. NFKB1 promoter/exon 1 luciferase reporter plasmid constructs containing the -94delATTG allele and transfected into either HeLa or HT-29 cell lines showed less promoter activity than comparable constructs containing the -94insATTG allele. Therefore, we have identified the first potentially functional polymorphism of NFKB1 and demonstrated its genetic association with a common human disease, ulcerative colitis.

Base Sequence↗

SUPFAM--a database of potential protein superfamily relationships derived by comparing sequence-based and structure-based families: implications for structural genomics and function annotation in genomes.

Members of a superfamily of proteins could result from divergent evolution of homologues with insignificant similarity in the amino acid sequences. A superfamily relationship is detected commonly after the three-dimensional structures of the proteins are determined using X-ray analysis or NMR. The SUPFAM database described here relates two homologous protein families in a multiple sequence alignment database of either known or unknown structure. The present release (1.1), which is the first version of the SUPFAM database, has been derived by analysing Pfam, which is one of the commonly used databases of multiple sequence alignments of homologous proteins. The first step in establishing SUPFAM is to relate Pfam families with the families in PALI, which is an alignment database of homologous proteins of known structure that is derived largely from SCOP. The second step involves relating Pfam families which could not be associated reliably with a protein superfamily of known structure. The profile matching procedure, IMPALA, has been used in these steps. The first step resulted in identification of 1280 Pfam families (out of 2697, i.e. 47%) which are related, either by close homologous connection to a SCOP family or by distant relationship to a SCOP family, potentially forming new superfamily connections. Using the profiles of 1417 Pfam families with apparently no structural information, an all-against-all comparison involving a sequence-profile match using IMPALA resulted in clustering of 67 homologous protein families of Pfam into 28 potential new superfamilies. Expansion of groups of related proteins of yet unknown structural information, as proposed in SUPFAM, should help in identifying 'priority proteins' for structure determination in structural genomics initiatives to expand the coverage of structural information in the protein sequence space. For example, we could assign 858 distinct Pfam domains in 2203 of the gene products in the genome of Mycobacterium tubercolosis. Fifty-one of these Pfam families of unknown structure could be clustered into 17 potentially new superfamilies forming good targets for structural genomics. SUPFAM database can be accessed at http://pauling.mbu.iisc.ernet.in/~supfam.

Animals↗

3D-GENOMICS: a database to compare structural and functional annotations of proteins between sequenced genomes.

The 3D-GENOMICS database (http://www.sbg.bio. ic.ac.uk/3dgenomics/) provides structural annotations for proteins from sequenced genomes. In August 2003 the database included data for 93 proteomes. The annotations stored in the database include homologous sequences from various sequence databases, domains from SCOP and Pfam, patterns from Prosite and other predicted sequence features such as transmembrane regions and coiled coils. In addition to annotations at the sequence level, several precomputed cross- proteome comparative analyses are available based on SCOP domain superfamily composition. Annotations are available to the user via a web interface to the database. Multiple points of entry are available so that a user is able to: (i) directly access annotations for a single protein sequence via keywords or accession codes, (ii) examine a sequence of interest chosen from a summary of annotations for a particular proteome, or (iii) access precomputed frequency-based cross-proteome comparative analyses.

Amino Acid Sequence↗

Analysis and functional annotation of an expressed sequence tag collection for tropical crop sugarcane.

To contribute to our understanding of the genome complexity of sugarcane, we undertook a large-scale expressed sequence tag (EST) program. More than 260,000 cDNA clones were partially sequenced from 26 standard cDNA libraries generated from different sugarcane tissues. After the processing of the sequences, 237,954 high-quality ESTs were identified. These ESTs were assembled into 43,141 putative transcripts. Of the assembled sequences, 35.6% presented no matches with existing sequences in public databases. A global analysis of the whole SUCEST data set indicated that 14,409 assembled sequences (33% of the total) contained at least one cDNA clone with a full-length insert. Annotation of the 43,141 assembled sequences associated almost 50% of the putative identified sugarcane genes with protein metabolism, cellular communication/signal transduction, bioenergetics, and stress responses. Inspection of the translated assembled sequences for conserved protein domains revealed 40,821 amino acid sequences with 1415 Pfam domains. Reassembling the consensus sequences of the 43,141 transcripts revealed a 22% redundancy in the first assembling. This indicated that possibly 33,620 unique genes had been identified and indicated that >90% of the sugarcane expressed genes were tagged.

Computational Biology↗

High-throughput functional affinity purification of mannose binding proteins from Oryza sativa.

We have used affinity chromatography in combination with mass spectrometry to isolate, identify, and assign a preliminary functional annotation to a large number of both known and novel proteins from rice. Rice (Oryza sativa) leaf, root, and seed tissue extracts were fractionated by column affinity chromatography using alpha-D-mannose as the ligand. Bound fractions were eluted and subjected to one-dimensional electrophoresis, followed by high-performance liquid chromatography-tandem mass spectrometric analysis of separated proteins. This multiplexed technology resulted in the isolation and identification of 136 distinct mannose binding proteins from rice. A comparative analysis demonstrates very little overlap of identified proteins between the respective tissues, and confirms the correctly compartmentalized presence of a significant number of proteins from largely tissue-specific biochemical pathways. Over 30% of the identified proteins with a previously annotated function are directly involved in sugar metabolism, including several highly expressed known rice lectins. Direct comparison of the peptide sequences identified in this study to those peptides identified in the most comprehensive survey of the rice proteome to date indicates that our current data represents a significant enrichment of proteins unique to this dataset. Nearly 15% of the identified proteins, identified on the basis of exact peptide matching to sequences in the rice genomic database, represent proteins without a previously known functional annotation, indicating the potential of this combined chromatographic approach to assign a preliminary function to novel proteins in a high-throughput fashion.

Binding, Competitive↗

Position-specific annotation of protein function based on multiple homologs.

I present in this work an algorithm for deriving protein functional annotations which are position-specific. The input is based on the results of a sequence similarity search of the query sequence against a sequence database. Strings of words are extracted from the descriptions of the proteins, and the correlation between proteins having the same descriptors and the amino acid conservation is used to compute a score that indicates which descriptor is likely to describe better the function of each particular residue. Analysis of the score curves and comparison of different functions allows an easy detection of parts of the sequence associated to different function. Different levels of functional specificity can be compared, allowing to choose the one that suits better the function of the protein. Immediate applications of this algorithm are, support for (automated) methods of protein functional annotation, and database coherence check.

Algorithms↗

RIKEN mouse genome encyclopedia.

We have been working to establish the comprehensive mouse full-length cDNA collection and sequence database to cover as many genes as we can, named Riken mouse genome encyclopedia. Recently we are constructing higher-level annotation (Functional ANnoTation Of Mouse cDNA; FANTOM) not only with homology search based annotation but also with expression data profile, mapping information and protein-protein database. More than 1,000,000 clones prepared from 163 tissues were end-sequenced to classify into 159,789 clusters and 60,770 representative clones were fully sequenced. As a conclusion, the 60,770 sequences contained 33,409 unique. The next generation of life science is clearly based on all of the genome information and resources. Based on our cDNA clones we developed the additional system to explore gene function. We developed cDNA microarray system to print all of these cDNA clones, protein-protein interaction screening system, protein-DNA interaction screening system and so on. The integrated database of all the information is very useful not only for analysis of gene transcriptional network and for the connection of gene to phenotype to facilitate positional candidate approach. In this talk, the prospect of the application of these genome resourced should be discussed. More information is available at the web page: http://genome.gsc.riken.go.jp/.

Animals↗

CDART: protein homology by domain architecture.

The Conserved Domain Architecture Retrieval Tool (CDART) performs similarity searches of the NCBI Entrez Protein Database based on domain architecture, defined as the sequential order of conserved domains in proteins. The algorithm finds protein similarities across significant evolutionary distances using sensitive protein domain profiles rather than by direct sequence similarity. Proteins similar to a query protein are grouped and scored by architecture. Relying on domain profiles allows CDART to be fast, and, because it relies on annotated functional domains, informative. Domain profiles are derived from several collections of domain definitions that include functional annotation. Searches can be further refined by taxonomy and by selecting domains of interest. CDART is available at http://www.ncbi.nlm.nih.gov/Structure/lexington/lexington.cgi.

BRCA1 Protein↗

Automatic annotation of protein function based on family identification.

Although genomes are being sequenced at an impressive rate, the information generated tells us little about protein function, which is slow to characterize by traditional methods. Automatic protein function annotation based on computational methods has alleviated this imbalance. The most powerful current approach for inferring the function of new proteins is by studying the annotations of their homologues, since their common origin is assumed to be reflected in their structure and function. Unfortunately, as proteins evolve they acquire new functions, so annotation based on homology must be carried out in the context of orthologues or subfamilies. Evolution adds new complications through domain shuffling: homology (or orthology) frequently corresponds to domains rather than complete proteins. Moreover, the function of a protein may be seen as the result of combining the functions of its domains. Additionally, automatic annotation has to deal with problems related to the annotations in the databases: errors (which are likely to be propagated), inconsistencies, or different degrees of function specification. We describe a method that addresses these difficulties for the annotation of protein function. Sequence relationships are detected and measured to obtain a map of the sequence space, which is searched for differentiated groups of proteins (similar to islands on the map), which are expected to have a common function and correspond to groups of orthologues or subfamilies. This mapmaking is done by applying a clustering algorithm based on Normalized cuts in graphs. The domain problem is addressed in a simple way: pairwise local alignments are analyzed to determine the extent to which they cover the entire sequence lengths of the two proteins. This analysis determines both what homologues are preferred for functional inheritance and the level of confidence of the annotation. To alleviate the problems associated with database annotations, the information on all the homologues that are grouped together with the query protein are taken into account to select the most representative functional descriptors. This method has been applied for the annotation of the genome of Buchnera aphidicola (specific host Baizongia pistaciae). Human inspection of the annotations allowed an estimation of accuracy of 94%; the different kinds of error that may appear when using this approach are described. Results can be accessed at http://www.pdg.cnb.uam.es/funcut.html. The programs are available upon request, although installation in other systems may be complicated.

Algorithms↗