Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,423 records · Page 79Linked to original sources

A functional approach to questions about life, death, and phosphorylation.

The success of the family of kinases as targets for small-molecule cancer therapeutics is probably best illustrated by the efficacy of the drug Gleevec. In spite of this, the function of many of the kinases in the mammalian genome remains unknown. In a recent paper, MacKeigan and colleagues report a functional genetic screen using RNA interference to identify kinases and phosphatases involved in programmed cell death (MacKeigan et al., 2005). Functional annotation is a prerequisite for selection of new drug targets. Such studies may therefore lay the foundation for the next generation of cancer drugs.

Antineoplastic Agents↗

ZCURVE: a new system for recognizing protein-coding genes in bacterial and archaeal genomes.

A new system, ZCURVE 1.0, for finding protein- coding genes in bacterial and archaeal genomes has been proposed. The current algorithm, which is based on the Z curve representation of the DNA sequences, lays stress on the global statistical features of protein-coding genes by taking the frequencies of bases at three codon positions into account. In ZCURVE 1.0, since only 33 parameters are used to characterize the coding sequences, it gives better consideration to both typical and atypical cases, whereas in Markov-model-based methods, e.g. Glimmer 2.02, thousands of parameters are trained, which may result in less adaptability. To compare the performance of the new system with that of Glimmer 2.02, both systems were run, respectively, for 18 genomes not annotated by the Glimmer system. Comparisons were also performed for predicting some function-known genes by both systems. Consequently, the average accuracy of both systems is well matched; however, ZCURVE 1.0 has more accurate gene start prediction, lower additional prediction rate and higher accuracy for the prediction of horizontally transferred genes. It is shown that the joint applications of both systems greatly improve gene-finding results. For a typical genome, e.g. Escherichia coli, the system ZCURVE 1.0 takes approximately 2 min on a Pentium III 866 PC without any human intervention. The system ZCURVE 1.0 is freely available at: http://tubic. tju.edu.cn/Zcurve_B/.

Algorithms↗

The Role of Small Segmental Duplications in Generating Identical Isoforms Through Alternative Splicing Sites.

Alternative splicing plays a crucial role in expanding proteomic diversity but can also generate identical isoforms under certain conditions. While mutually exclusive splicing of tandem exons has occasionally been reported to produce identical isoforms, the extent to which other splicing events contribute to this phenomenon remains unclear. In this study, we demonstrate that alternative 5' and 3' splice site selection can also lead to the formation of identical isoforms, providing an additional type of splicing event for functional redundancy in transcriptomes. To address this, we analyzed reference genome annotations from 15 plant species, including Arabidopsis thaliana and wheat (Triticum aestivum), obtained from the RefSeq database. Identical isoforms were computationally defined as transcripts with distinct exon-intron structures but identical coding sequences. Our analysis reveals that the majority of alternative 5' and 3' fragments originate from small segmental duplications, suggesting that sequence repetition within gene regions facilitates the emergence of such splicing patterns. We also observed differences in the annotated 5' UTRs of some identical isoforms. However, since the alternative splicing sites themselves were not located within UTRs, these differences may reflect annotation uncertainty rather than genuine AS-derived variation. Given that UTR predictions in reference databases are not always precise, such observations should be interpreted cautiously. Expression analysis using an isoform-specific k-mer approach confirmed that identical isoforms can be differentially regulated. These findings suggest that, beyond expanding protein diversity, alternative splicing can also generate redundant isoforms that are differentially expressed at the RNA level, indicating potential regulatory roles. By elucidating the structural and regulatory factors contributing to the formation and retention of identical isoforms, our study provides new insights into the evolutionary and functional significance of alternative splicing in plants.

Alternative Splicing↗

TCDB: the Transporter Classification Database for membrane transport protein analyses and information.

The Transporter Classification Database (TCDB) is a web accessible, curated, relational database containing sequence, classification, structural, functional and evolutionary information about transport systems from a variety of living organisms. TCDB is a curated repository for factual information compiled from >10,000 references, encompassing approximately 3000 representative transporters and putative transporters, classified into >400 families. The transporter classification (TC) system is an International Union of Biochemistry and Molecular Biology approved system of nomenclature for transport protein classification. TCDB is freely accessible at http://www.tcdb.org. The web interface provides several different methods for accessing the data, including step-by-step access to hierarchical classification, direct search by sequence or TC number and full-text searching. The functional ontology that underlies the database structure facilitates powerful query searches that yield valuable data in a quick and easy way. The TCDB website also offers several tools specifically designed for analyzing the unique characteristics of transport proteins. TCDB not only provides curated information and a tool for classifying newly identified membrane proteins, but also serves as a genome transporter-annotation tool.

Databases, Protein↗

The predicted candidates of Arabidopsis plastid inner envelope membrane proteins and their expression profiles.

Plastid envelope proteins from the Arabidopsis nuclear genome were predicted using computational methods. Selection criteria were: first, to find proteins with NH(2)-terminal plastid-targeting peptides from all annotated open reading frames from Arabidopsis; second, to search for proteins with membrane-spanning domains among the predicted plastidial-targeted proteins; and third, to subtract known thylakoid membrane proteins. Five hundred forty-one proteins were selected as potential candidates of the Arabidopsis plastid inner envelope membrane proteins (AtPEM candidates). Only 34% (183) of the AtPEM candidates could be assigned to putative functions based on sequence similarity to proteins of known function (compared with the 69% function assignment of the total predicted proteins in the genome). Of the 183 candidates with assigned functions, 40% were classified in the category of "transport facilitation," indicating that this collection is highly enriched in membrane transporters. Information on the predicted proteins, tissue expression data from expressed sequence tags and microarrays, and publicly available T-DNA insertion lines were collected. The data set complements proteomic-based efforts in the increased detection of integral membrane proteins, low-abundance proteins, or those not expressed in tissues selected for proteomic analysis. Digital northern analysis of expressed sequence tags suggested that the transcript levels of most AtPEM candidates were relatively constant among different tissues in contrast to stroma and the thylakoid proteins. However, both digital northern and microarray analyses identified a number of AtPEM candidates with tissue-specific expression patterns.

Arabidopsis↗

Cellular response of Shewanella oneidensis to strontium stress.

The physiology and transcriptome dynamics of the metal ion-reducing bacterium Shewanella oneidensis strain MR-1 in response to nonradioactive strontium (Sr) exposure were investigated. Studies indicated that MR-1 was able to grow aerobically in complex medium in the presence of 180 mM SrCl2 but showed severe growth inhibition at levels above that concentration. Temporal gene expression profiles were generated from aerobically grown, mid-exponential-phase MR-1 cells shocked with 180 mM SrCl2 and analyzed for significant differences in mRNA abundance with reference to data for nonstressed MR-1 cells. Genes with annotated functions in siderophore biosynthesis and iron transport were among the most highly induced (>100-fold [P < 0.05]) open reading frames in response to acute Sr stress, and a mutant (SO3032::pKNOCK) defective in siderophore production was found to be hypersensitive to SrCl2 exposure, compared to parental and wild-type strains. Transcripts encoding multidrug and heavy metal efflux pumps, proteins involved in osmotic adaptation, sulfate ABC transporters, and assimilative sulfur metabolism enzymes also were differentially expressed following Sr exposure but at levels that were several orders of magnitude lower than those for iron transport genes. Precipitate formation was observed during aerobic growth of MR-1 in broth cultures amended with 50, 100, or 150 mM SrCl2 but not in cultures of the SO3032::pKNOCK mutant or in the abiotic control. Chemical analysis of this precipitate using laser-induced breakdown spectroscopy and static secondary ion mass spectrometry indicated extracellular solid-phase sequestration of Sr, with at least a portion of the heavy metal associated with carbonate phases.

Bacterial Proteins↗

Annotation and analysis of 10,000 expressed sequence tags from developing mouse eye and adult retina.

BACKGROUND: As a biomarker of cellular activities, the transcriptome of a specific tissue or cell type during development and disease is of great biomedical interest. We have generated and analyzed 10,000 expressed sequence tags (ESTs) from three mouse eye tissue cDNA libraries: embryonic day 15.5 (M15E) eye, postnatal day 2 (M2PN) eye and adult retina (MRA). RESULTS: Annotation of 8,633 non-mitochondrial and non-ribosomal high-quality ESTs revealed that 57% of the sequences represent known genes and 43% are unknown or novel ESTs, with M15E having the highest percentage of novel ESTs. Of these, 2,361 ESTs correspond to 747 unique genes and the remaining 6,272 are represented only once. Phototransduction genes are preferentially identified in MRA, whereas transcripts for cell structure and regulatory proteins are highly expressed in the developing eye. Map locations of human orthologs of known genes uncovered a high density of ocular genes on chromosome 17, and identified 277 genes in the critical regions of 37 retinal disease loci. In silico expression profiling identified 210 genes and/or ESTs over-expressed in the eye; of these, more than 26 are known to have vital retinal function. Comparisons between libraries provided a list of temporally regulated genes and/or ESTs. A few of these were validated by qRT-PCR analysis. CONCLUSIONS: Our studies present a large number of potentially interesting genes for biological investigation, and the annotated EST set provides a useful resource for microarray and functional genomic studies.

Aging↗

Comparative accuracy of methods for protein sequence similarity search.

MOTIVATION: Searching a protein sequence database for homologs is a powerful tool for discovering the structure and function of a sequence. Two new methods for searching sequence databases have recently been described: Probabilistic Smith-Waterman (PSW), which is based on Hidden Markov models for a single sequence using a standard scoring matrix, and a new version of BLAST (WU-BLAST2), which uses Sum statistics for gapped alignments. RESULTS: This paper compares and contrasts the effectiveness of these methods with three older methods (Smith-Waterman: SSEARCH, FASTA and BLASTP). The analysis indicates that the new methods are useful, and often offer improved accuracy. These tools are compared using a curated (by Bill Pearson) version of the annotated portion of PIR 39. Three different statistical criteria are utilized: equivalence number, minimum errors and the receiver operating characteristic. For complete-length protein query sequences from large families, PSW's accuracy is superior to that of the other methods, but its accuracy is poor when used with partial-length query sequences. False negatives are twice as common as false positives irrespective of the search methods if a family-specific threshold score that minimizes the total number of errors (i.e. the most favorable threshold score possible) is used. Thus, sensitivity, not selectivity, is the major problem. Among the analyzed methods using default parameters, the best accuracy was obtained from SSEARCH and PSW for complete-length proteins, and the two BLAST programs, plus SSEARCH, for partial-length proteins.

Databases, Factual↗

The complete and annotated mitochondrial genome of Hemileia vastatrix Race I, causal agent of coffee leaf rust.

Hemileia vastatrix is the fungal pathogen responsible for coffee leaf rust (CLR), the most economically important disease of Coffea arabica worldwide. Recently, the nuclear genome of this fungus was completely deciphered. However, the mitochondrial genome of H. vastatrix has remained undercharacterized. Here, we present the complete, circularized mitochondrial genome of H. vastatrix Race I (isolate HvRI), assembled using a hybrid approach combining PacBio HiFi long reads and BGIseq short reads. The genome is 173,525&#xa0;bp in length with a GC content of 33.1% and encodes 41 functional genes, including 15 protein-coding genes, 2 rRNAs, and 24 tRNAs. The assembly reveals significant structural complexity, driven by intron expansion in the cox1 and cob genes. Notably, the atp8 gene contains a group II intron, rare for this locus, whose internal open reading frame displays evidence of pseudogenization via internal stop codons.. We also characterized a putative replication initiation zone (~1.2&#xa0;kb) defined by a poly-G homopolymer and conserved regulatory motifs. The mitogenome of the HvRI isolate does not contain cob mutations that lead to amino acid substitutions G143A and F129L associated with the quinone outside inhibitor (QoI) fungicide resistance. This high-quality mitogenome is an important resource for comparative mitogenomics, population diversity studies, and the molecular surveillance of QoI fungicide resistance.

Genome, Mitochondrial↗

Unusual usage of AGG and TTG codons in humans and their viruses.

Prior analysis on human protein-coding DNA sequences has identified local base composition as the primary predictor of synonymous codon usage. However, in many organisms, codon usage is influenced by natural selection, particularly for efficient expression of functional gene products. Because viruses are expected to evolve codon usage in the context of their host's molecular machinery, their genomes provide another window into the forces that guide their host's molecular evolution. Factor analysis was performed on codon usage of 16,654 genes annotated in Build 34 of the human genome, and the primary factor was correlated strongly with local base composition. However, two codons, AGG and TTG, rose in frequency as all other C- and G-ending codons decreased in frequency. These two codons were the only C- or G-ending codons with usages that negatively correlated with gene expression. Variation among viruses in codon usage also strongly reflects variation in base composition and, again, AGG and TTG decrease in frequency as all other C- and G-ending codons increase in frequency. It appears that usages of these two codons can not be explained by local compositional biases, implying a more direct role of natural selection on codon usage in humans.

Amino Acids↗

RNAs everywhere: genome-wide annotation of structured RNAs.

Starting with the discovery of microRNAs and the advent of genome-wide transcriptomics, non-protein-coding transcripts have moved from a fringe topic to a central field research in molecular biology. In this contribution we review the state of the art of "computational RNomics", i.e., the bioinformatics approaches to genome-wide RNA annotation. Instead of rehashing results from recently published surveys in detail, we focus here on the open problem in the field, namely (functional) annotation of the plethora of putative RNAs. A series of exploratory studies are used to provide non-trivial examples for the discussion of some of the difficulties.

Biological Evolution↗

Computational approaches to protein-protein interaction.

The interactions between proteins allow the cell's life. A number of experimental, genome-wide, high-throughput studies have been devoted to the determination of protein-protein interactions and the consequent interaction networks. Here, the bioinformatics methods dealing with protein-protein interactions and interaction network are overviewed. 1. Interaction databases developed to collect and annotate this immense amount of data; 2. Automated data mining techniques developed to extract information about interactions from the published literature; 3. Computational methods to assess the experimental results developed as a consequence of the finding that the results of high-throughput methods are rather inaccurate; 4. Exploitation of the information provided by protein interaction networks in order to predict functional features of the proteins; and 5. Prediction of protein-protein interactions.

Algorithms↗

Making connections between novel transcription factors and their DNA motifs.

The key components of a transcriptional regulatory network are the connections between trans-acting transcription factors and cis-acting DNA-binding sites. In spite of several decades of intense research, only a fraction of the estimated approximately 300 transcription factors in Escherichia coli have been linked to some of their binding sites in the genome. In this paper, we present a computational method to connect novel transcription factors and DNA motifs in E. coli. Our method uses three types of mutually independent information, two of which are gleaned by comparative analysis of multiple genomes and the third one derived from similarities of transcription-factor-DNA-binding-site interactions. The different types of information are combined to calculate the probability of a given transcription-factor-DNA-motif pair being a true pair. Tested on a study set of transcription factors and their DNA motifs, our method has a prediction accuracy of 59% for the top predictions and 85% for the top three predictions. When applied to 99 novel transcription factors and 70 novel DNA motifs, our method predicted 64 transcription-factor-DNA-motif pairs. Supporting evidence for some of the predicted pairs is presented. Functional annotations are made for 23 novel transcription factors based on the predicted transcription-factor-DNA-motif connections.

Algorithms↗

A novel genetic island of meningitic Escherichia coli K1 containing the ibeA invasion gene (GimA): functional annotation and carbon-source-regulated invasion of human brain microvascular endothelial cells.

The IbeA (ibe10) gene is an invasion determinant contributing to E. coli K1 invasion of the blood-brain barrier. This gene has been cloned and characterized from the chromosome of an invasive cerebrospinal fluid isolate of E. coli K1, strain RS218 (018:K1: H7). In the present study, a genetic island of meningitic E. coli containing ibeA (GimA) has been identified. A 20.3-kb genomic DNA island unique to E. coli K1 strains has been cloned and sequenced from an RS218 E. coli K1 genomic DNA library. Fourteen new genes have been identified in addition to the ibeA. The DNA sequence analysis indicated that the ibeA gene cluster was localized to the 98 min region and consisted of four operons, ptnIPKC, cglDTEC, gcxKRCI and ibeRAT. The G+C content (46.2%) of unique regions of the island is substantially different from that (50.8%) of the rest of the E. coli chromosome. By computer-assisted analysis of the sequences with DNA and protein databases (GenBank and PROSITE databases), the functions of the gene products could be anticipated, and were assigned to the functional categories of proteins relating to carbon source metabolism and substrate transportation. Glucose was shown to enhance E. coli penetration of human brain microvascular endothelial cells and exogenous cAMP was able to block the stimulating effect of glucose, suggesting that catabolic regulation may play a role in control of E. coli K1 invasion gene expression. Our data suggest that this genetic island may contribute to E. coli invasion of the blood-brain barrier through a carbon-source-regulated process.

Amino Acid Sequence↗

Biological function of unannotated transcription during the early development of Drosophila melanogaster.

Many animal and plant genomes are transcribed much more extensively than current annotations predict. However, the biological function of these unannotated transcribed regions is largely unknown. Approximately 7% and 23% of the detected transcribed nucleotides during D. melanogaster embryogenesis map to unannotated intergenic and intronic regions, respectively. Based on computational analysis of coordinated transcription, we conservatively estimate that 29% of all unannotated transcribed sequences function as missed or alternative exons of well-characterized protein-coding genes. We estimate that 15.6% of intergenic transcribed regions function as missed or alternative transcription start sites (TSS) used by 11.4% of the expressed protein-coding genes. Identification of P element mutations within or near newly identified 5' exons provides a strategy for mapping previously uncharacterized mutations to their respective genes. Collectively, these data indicate that at least 85% of the fly genome is transcribed and processed into mature transcripts representing at least 30% of the fly genome.

Amino Acid Sequence↗

GARBAN: genomic analysis and rapid biological annotation of cDNA microarray and proteomic data.

SUMMARY: Genomic Analysis and Rapid Biological ANnotation (GARBAN) is a new tool that provides an integrated framework to analyze simultaneously and compare multiple data sets derived from microarray or proteomic experiments. It carries out automated classifications of genes or proteins according to the criteria of the Gene Ontology Consortium at a level of depth defined by the user. Additionally, it performs clustering analysis of all sets based on functional categories or on differential expression levels. GARBAN also provides graphical representations of the biological pathways in which all the genes/proteins participate. AVAILABILITY: http://garban.tecnun.es.

Algorithms↗

DOCKGROUND resource for studying protein-protein interfaces.

MOTIVATION: Public resources for studying protein interfaces are necessary for better understanding of molecular recognition and developing intermolecular potentials, search procedures and scoring functions for the prediction of protein complexes. RESULTS: The first release of the DOCKGROUND resource implements a comprehensive database of co-crystallized (bound-bound) protein-protein complexes, providing foundation for the upcoming expansion to unbound (experimental and simulated) protein-protein complexes, modeled protein-protein complexes and systematic sets of docking decoys. The bound-bound part of DOCKGROUND is a relational database of annotated structures based on the Biological Unit file (Biounit) provided by the RCSB as a separated file containing probable biological molecule. DOCKGROUND is automatically updated to reflect the growth of PDB. It contains 67,220 pairwise complexes that rely on 14,913 Biounit entries from 34,778 PDB entries (January 30, 2006). The database includes a dynamic generation of non-redundant datasets of pairwise complexes based either on the structural similarity (SCOP classification) or on user-defined sequence identity. The growing DOCKGROUND resource is designed to become a comprehensive public environment for developing and validating new methodologies for modeling of protein interactions. AVAILABILITY: DOCKGROUND is available at http://dockground.bioinformatics.ku.edu. The current first release implements the bound-bound part.

Binding Sites↗

Genomics and proteomics of bone cancer.

Although the control of bone metastasis has been the focus of intensive investigation, relatively little is known about the molecular mechanisms that regulate or predict the process, even though widespread skeletal dissemination is an important step in the progression of many tumors. As a result, understanding the complex interactions contributing to the metastatic behavior of tumor cells is essential for the development of effective therapies. Using a state-of-the-art combination of gene expression profiling and functional annotation of human tumor cells, and surface-enhanced laser desorption/ionization time-of-flight mass spectrometry of patient serum, we have shown that changes in tumor biochemistry correlate with disease progression and help to define the aggressive tumor phenotype. Based on these approaches, it is apparent that the metastatic phenotype of tumor cells is extremely complex. The identification of the phenotype of tumor cells has benefited greatly from the application of gene expression profiling (microarray analysis). This technology has been used by many investigators to identify changes in gene expression and cytokine and growth factor elaboration (such as interleukin 8). The tumor phenotype(s) presumably also include changes in the cell surface carbohydrate profile (via altered glycosyltransferase expression) and heparan sulfate expression (via increased heparanase activity), to name but a few. These specific alterations in gene expression, identified by functional annotation of accumulated microarray data, have been validated using a variety of approaches. Collectively, the data described here suggest that each of these activities is associated with distinct aspects of the aggressive tumor cell phenotype. Collectively, the data suggest that multiple factors constitute the complex phenotype of metastatic tumor cells. In particular, the differences observed in gene expression profiles and serum protein biomarkers play a critical role in defining the mechanisms responsible for bone-specific colonization and growth of tumors in bone. Future studies will identify the mechanisms that participate in the formation of secondary tumor growths of cancers in bone.

Biomarkers, Tumor↗