Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,225 records · Page 68Linked to original sources

Central functions of the lumenal and peripheral thylakoid proteome of Arabidopsis determined by experimentation and genome-wide prediction.

Experimental proteome analysis was combined with a genome-wide prediction screen to characterize the protein content of the thylakoid lumen of Arabidopsis chloroplasts. Soluble thylakoid proteins were separated by two-dimensional electrophoresis and identified by mass spectrometry. The identities of 81 proteins were established, and N termini were sequenced to validate localization prediction. Gene annotation of the identified proteins was corrected by experimental data, and an interesting case of alternative splicing was discovered. Expression of a surprising number of paralogs was detected. Expression of five isomerases of different classes suggests strong (un)folding activity in the thylakoid lumen. These isomerases possibly are connected to a network of peripheral and lumenal proteins involved in antioxidative response, including peroxiredoxins, m-type thioredoxins, and a lumenal ascorbate peroxidase. Characteristics of the experimentally identified lumenal proteins and their orthologs were used for a genome-wide prediction of the lumenal proteome. Lumenal proteins with a typical twin-arginine translocation motif were predicted with good accuracy and sensitivity and included additional isomerases and proteases. Thus, prime functions of the lumenal proteome include assistance in the folding and proteolysis of thylakoid proteins as well as protection against oxidative stress. Many of the predicted lumenal proteins must be present at concentrations at least 10,000-fold lower than proteins of the photosynthetic apparatus.

Alternative Splicing↗

Characterization of long cDNA clones from human adult spleen. II. The complete sequences of 81 cDNA clones.

To accumulate information on the coding sequences (CDSs) of unidentified genes, we have conducted a sequencing project of human long cDNA clones. Both the end sequences of approximately 10,000 cDNA clones from two size-fractionated human spleen cDNA libraries (average sizes of 4.5 kb and 5.6 kb) were determined by single-pass sequencing to select cDNAs with unidentified sequences. We herein present the entire sequences of 81 cDNA clones, most of which were selected by two approaches based on their protein-coding potentialities in silico: Fifty-eight cDNA clones were selected as those having protein-coding potentialities at the 5'-end of single-pass sequences by applying the GeneMark analysis; and 20 cDNA clones were selected as those expected to encode proteins larger than 100 amino acid residues by analysis of the human genome sequences flanked by both the end sequences of cDNAs using the GENSCAN gene prediction program. In addition to these newly identified cDNAs, three cDNA clones were isolated by colony hybridization experiments using probes corresponding to known gene sequences since these cDNAs are likely to contain considerable amounts of new information regarding the genes already annotated. The sequence data indicated that the average sizes of the inserts and corresponding CDSs of cDNA clones analyzed here were 5.0 kb and 2.0 kb (670 amino acid residues), respectively. From the results of homology and motif searches against the public databases, functional categories of the 29 predicted gene products could be assigned; 86% of these predicted gene products (25 gene products) were classified into proteins relating to cell signaling/communication, nucleic acid management, and cell structure/motility.

Adult↗

Automatic detection of subsystem/pathway variants in genome analysis.

MOTIVATION: Proteins work together in pathways and networks, collectively comprising the cellular machinery. A subsystem (a generalization of pathway concept) is a group of related functional roles (such as enzymes) jointly involved in a specific aspect of the cellular machinery. Subsystems provide a natural framework for comparative genome analysis and functional annotation. A subsystem may be implemented in a number of different functional variants in individual species. In order to reliably project functional assignments across multiple genomes, we have to be able to identify the variants implemented in each genome. The analysis of such variants across diverse species is an interesting problem by itself and may provide new evolutionary insights. However, no computational techniques are presently available for an automated detection and analysis of subsystem variants. RESULTS: Here we formulate the subsystem variant detection problem as finding the minimum number of subgraphs of a subsystem, which is represented as a graph, and solve the optimization problem by integer programming approach. The performance of our method was tested on subsystems encoded in the SEED, a genomic integration platform developed by the Fellowship for Interpretation of Genomes as a component of a large-scale effort on comparative analysis and annotation of multiple diverse genomes. Here we illustrate the results obtained for two expert-encoded subsystems of the biosynthesis of Coenzyme A and FMN/FAD cofactors. Applications of variant detection, to support genomic annotations and to assess divergence of species, are briefly discussed in the context of these universally conserved and essential metabolic subsystems. SUPPLEMENTARY INFORMATION: The details of the variant detection results are available at http://ffas.burnham.org/svar/supp.html.

Animals↗

BIOVERSE: enhancements to the framework for structural, functional and contextual modeling of proteins and proteomes.

We have made a number of enhancements to the previously described Bioverse web server and computational biology framework (http://bioverse.compbio.washington.edu). In this update, we provide an overview of the new features available that include: (i) expansion of the number of organisms represented in the Bioverse and addition of new data sources and novel prediction techniques not available elsewhere, including network-based annotation; (ii) reengineering the database backend and supporting code resulting in significant speed, search and ease-of use improvements; and (iii) creation of a stateful and dynamic web application frontend to improve interface speed and usability. Integrated Java-based applications also allow dynamic visualization of real and predicted protein interaction networks.

Computer Graphics↗

Proteome analysis of Neisseria meningitidis serogroup A.

Neisseria meningitidis is an encapsulated Gram-negative bacterium responsible for significant morbidity and mortality worldwide. Meningococci are opportunistic pathogens, carried in the nasopharynx of approximately 10% of asymptomatic adults. Occasionally they enter the bloodstream to cause septicaemia and meningitis. Meningococci are classified into serogroups on the basis of polysaccharide capsule diversity, and serogroup A strains have caused major epidemics mainly in the developing world. Here we describe a two-dimensional gel electrophoresis protein map of the serogroup A strain Z4970, a clinical isolate classified as ancestral to several pandemic waves. To our knowledge this is the first systematically annotated proteomic map for N. meningitidis. Total protein samples from bacteria grown on GC-agar were electrophoretically separated and protein species were identified by matrix-assisted laser desorption/ionization time of flight spectrometry. We identified the products of 273 genes, covering several functional classes, including 94 proteins so far considered as hypothetical. We also describe several protein species encoded by genes reported by DNA microarray studies as being regulated in physiological conditions which are relevant to natural meningococcal pathogenicity. Since menA differs from other serogroups by having a fairly stable clonal population structure (i.e. with a low degree of variability), we envisaged comparative mapping as a useful tool for microevolution studies, in conjunction with established genotyping methods. As a proof of principle, we performed a comparative analysis on the B subunit of the meningococcal transferrin receptor, a vaccine candidate encoded by the tbpB gene, and a known marker of population diversity in meningococci. The results show that TbpB spot pattern variation observed in the maps of nine clinical isolates from diverse epidemic spreads, fits previous analyses based on allelic variations of the tbpB gene.

Agar↗

Gene expression and EST analyses of Ustilago maydis germinating teliospores.

Ustilago maydis grows in its host Zea mays eliciting the formation of obvious tumors that are full of black teliospores. Teliospores are thick-walled, dormant, diploid cells that have evolved for dispersal and survival of the pathogen. Their germination leads to new rounds of infection and is temporally linked to meiosis. We are investigating gene expression during teliospore germination to gain insight into the control of this process. Here we identify genes expressed through creation of an expressed sequence tag (EST) library. We generated 2871 ESTs that are assembled into 1293 contiguous sequences. Based upon a blast search similarity cutoff of E < or =10(-5) 38% of all contigs were orphans while 62% showed similarity to sequences in the protein database. Analyses of blast searches were used to functionally classify genes. Northern hybridizations using specific cDNA clones reveal a relative level of expression consistent with the number of sequences per contig. Identified genes and expression information provide a base for genome annotation of U. maydis and further investigation of teliospore germination and pathogenesis.

Blotting, Northern↗

Identification of novel tumour-associated genes differentially expressed in the process of squamous cell cancer development.

Chemically induced mouse skin carcinogenesis represents the most extensively utilized animal model to unravel the multistage nature of tumour development and to design novel therapeutic concepts of human epithelial neoplasia. We combined this tumour model with comprehensive gene expression analysis and could identify a large set of novel tumour-associated genes that have not been associated with epithelial skin cancer development yet. Expression data of selected genes were confirmed by semiquantitative and quantitative RT-PCR as well as in situ hybridization and immunofluorescence analysis on mouse tumour sections. Enhanced expression of genes identified in our screen was also demonstrated in mouse keratinocyte cell lines that form tumours in vivo. Self-organizing map clustering was performed to identify different kinetics of gene expression and coregulation during skin cancer progression. Detailed analysis of differential expressed genes according to their functional annotation confirmed the involvement of several biological processes, such as regulation of cell cycle, apoptosis, extracellular proteolysis and cell adhesion, during skin malignancy. Finally, we detected high transcript levels of ANXA1, LCN2 and S100A8 as well as reduced levels for NDR2 protein in human skin tumour specimens demonstrating that tumour-associated genes identified in the chemically induced tumour model might be of great relevance for the understanding of human epithelial malignancies as well.

Animals↗

Data integration and visualization system for enabling conceptual biology.

MOTIVATION: Integration of heterogeneous data in life sciences is a growing and recognized challenge. The problem is not only to enable the study of such data within the context of a biological question but also more fundamentally, how to represent the available knowledge and make it accessible for mining. RESULTS: Our integration approach is based on the premise that relationships between biological entities can be represented as a complex network. The context dependency is achieved by a judicious use of distance measures on these networks. The biological entities and the distances between them are mapped for the purpose of visualization into the lower dimensional space using the Sammon's mapping. The system implementation is based on a multi-tier architecture using a native XML database and a software tool for querying and visualizing complex biological networks. The functionality of our system is demonstrated with two examples: (1) A multiple pathway retrieval, in which, given a pathway name, the system finds all the relationships related to the query by checking available metabolic pathway, transcriptional, signaling, protein-protein interaction and ontology annotation resources and (2) A protein neighborhood search, in which given a protein name, the system finds all its connected entities within a specified depth. These two examples show that our system is able to conceptually traverse different databases to produce testable hypotheses and lead towards answers to complex biological questions.

Computational Biology↗

Systematic identification of pseudogenes through whole genome expression evidence profiling.

The identification of pseudogenes is an integral and significant part of the genome annotation because of their abundance and their impact on the experimental analysis of functional genes. Most of the computational annotation systems are not optimized for systematic pseudogene recognition, often annotating pseudogenes as functional genes, and users then propagate these errors to subsequent analyses and interpretations. In order to validate gene annotations and to identify pseudogenes that are potentially mis-annotated, we developed a novel approach based on whole genome profiling of existing transcript and protein sequences. This method has two important features: (i) equally detects both processed and non-processed pseudogenes and (ii) can identify transcribed pseudogenes. Applying this method to the human Ensembl gene predictions, we discovered that 2011 (9% of total) Ensembl genes in the categories of known and novel might be pseudogenes based on expression evidence. Of these, 1200 genes are found to have no existing evidence of transcription, and 811 genes are found with transcription evidence but contain significant translation disruption. Approximately 40% of the 2011 identified pseudogenes presented a multi-exon structure, representing non-processed pseudogenes. We have demonstrated the power of whole genome profiling of expression sequences to improve the accuracy of gene annotations.

Computational Biology↗

Complete genome sequence of the marine, chemolithoautotrophic, ammonia-oxidizing bacterium Nitrosococcus oceani ATCC 19707.

The gammaproteobacterium Nitrosococcus oceani (ATCC 19707) is a gram-negative obligate chemolithoautotroph capable of extracting energy and reducing power from the oxidation of ammonia to nitrite. Sequencing and annotation of the genome revealed a single circular chromosome (3,481,691 bp; G+C content of 50.4%) and a plasmid (40,420 bp) that contain 3,052 and 41 candidate protein-encoding genes, respectively. The genes encoding proteins necessary for the function of known modes of lithotrophy and autotrophy were identified. Contrary to betaproteobacterial nitrifier genomes, the N. oceani genome contained two complete rrn operons. In contrast, only one copy of the genes needed to synthesize functional ammonia monooxygenase and hydroxylamine oxidoreductase, as well as the proteins that relay the extracted electrons to a terminal electron acceptor, were identified. The N. oceani genome contained genes for 13 complete two-component systems. The genome also contained all the genes needed to reconstruct complete central pathways, the tricarboxylic acid cycle, and the Embden-Meyerhof-Parnass and pentose phosphate pathways. The N. oceani genome contains the genes required to store and utilize energy from glycogen inclusion bodies and sucrose. Polyphosphate and pyrophosphate appear to be integrated in this bacterium's energy metabolism, stress tolerance, and ability to assimilate carbon via gluconeogenesis. One set of genes for type I ribulose-1,5-bisphosphate carboxylase/oxygenase was identified, while genes necessary for methanotrophy and for carboxysome formation were not identified. The N. oceani genome contains two copies each of the genes or operons necessary to assemble functional complexes I and IV as well as ATP synthase (one H(+)-dependent F(0)F(1) type, one Na(+)-dependent V type).

Adenosine Triphosphate↗

ATIDB: Arabidopsis thaliana insertion database.

Insertional mutagenesis techniques, including transposon- and T-DNA-mediated mutagenesis, are key resources for systematic identification of gene function in the model plant species Arabidopsis thaliana. We have developed a database (http://atidb.cshl.org/) for archiving, searching and analyzing insertional mutagenesis lines. Flanking sequences from approximately 10 500 insertion lines (including transposon and T-DNA insertions) from several tagging programs in Arabidopsis were mapped to the genome sequence through our annotation system before being entered into the database. The database front end provides World Wide Web searching and analyzing interfaces for genome researchers and other biologists. Users can search the database to identify insertions in a particular gene or perform genome-wide analysis to study the distribution and preference of insertions. Tools integrated with the database include a graphical genome browser, a protein search function, a graphical representation of the insertion distribution and a Blast search function. The database is based on open source components and is available under an open source license.

Amino Acid Sequence↗

Improved diagnosis of patients with rare diseases through the application of constrained coding region annotation and de novo status.

PURPOSE: Identifying the pathogenic variant in a patient with rare disease (RD) is the first step in ending their diagnostic odyssey. De novo (Dn) variants affecting protein-coding DNA are a well-established cause of Mendelian disorders in patients with RD. Constrained coding regions (CCRs) are specific segments of coding DNA that are devoid of functional variants in healthy individuals. METHODS: We evaluated the diagnostic utility of incorporating combined Dn/CCR status into the variant prioritization cascade for patients with RD that have undergone genomic sequencing. Using the Genomics England 100,000 Genomes Project v12, we selected 3090 trios that have undergone diagnostic evaluation and been analyzed with an advanced Dn identification pipeline. RESULTS: Our analysis shows that the diagnostic rate increased from 71% in the full cohort to 87% for Dn/CCR variants. Of note, manual evaluation of the Dn/CCR variants from undiagnosed patients with clinical follow-up revealed a diagnosis for 13 further patients. This outcome increases the diagnostic rate for Dn/CCR variants to 91% and suggests that the application of this metric can prioritize diagnostic variants in undiagnosed patients. CONCLUSION: We demonstrate the potential clinical utility of performing bespoke Dn analyses of patients with RD and for incorporating CCR information into the filtering cascade to prioritize pathogenic variants.

Humans↗

Computationally analyzing the possible biological function of YJL103C--an ORF potentially involved in the regulation of energy process in yeast.

Although the complete genomes of a number of organisms have been sequenced, the biological functions of many genes are still not known. Because experimentally studying the functions of those genes one by one requires tremendous time, it is vital to use published resources like microarray gene expression data for computational analysis of gene functions. One example is YJL103C, a yeast gene of unknown function in the Saccharomyces Genome Database (SGD). It is possible to quickly infer its biological function by computational analysis. In this study, we present an efficient model to explore the biological function of a novel gene using microarray data. We showed that the expression pattern of YJL103C is most similar to the genes in the energy group and respiratory chain subgroup. We further found that YJL103C contains a HAP2,3,4 box in its promoter region and a cytochrome C heme-binding signature in its protein sequence. Our findings define a potential role for YJL103C in the regulation of energy metabolism, specifically in the process of oxidative phosphorylation. Similar bioinformatics methods can be applied to infer the biological functions of other novel genes in organisms for which microarray data are available. In this work, we selected a single gene of unknown function as a case study. By focusing on the power of computer analysis and bioinformatics on the available microarray data, we have determined the likely biological function of YJL103C. Our study provides a method by which to explore the potential function of other genes currently annotated as having an unknown function in any organism for which global gene expression data are available.

Binding Sites↗

Rapid creation of BAC-based human artificial chromosome vectors by transposition with synthetic alpha-satellite arrays.

Efficient construction of BAC-based human artificial chromosomes (HACs) requires optimization of each key functional unit as well as development of techniques for the rapid and reliable manipulation of high-molecular weight BAC vectors. Here, we have created synthetic chromosome 17-derived alpha-satellite arrays, based on the 16-monomer repeat length typical of natural D17Z1 arrays, in which the consensus CENP-B box elements are either completely absent (0/16 monomers) or increased in density (16/16 monomers) compared to D17Z1 alpha-satellite (5/16 monomers). Using these vectors, we show that the presence of CENP-B box elements is a requirement for efficient de novo centromere formation and that increasing the density of CENP-B box elements may enhance the efficiency of de novo centromere formation. Furthermore, we have developed a novel, high-throughput methodology that permits the rapid conversion of any genomic BAC target into a HAC vector by transposon-mediated modification with synthetic alpha-satellite arrays and other key functional units. Taken together, these approaches offer the potential to significantly advance the utility of BAC-based HACs for functional annotation of the genome and for applications in gene transfer.

Autoantigens↗

Identification of the R2R3-MYB gene family in wild jujube (Ziziphus jujuba var. spinosa) and analysis of its expression under drought stress.

BACKGROUND: R2R3-MYB gene family serves as a pivotal regulatory factor in plant growth, development, and responses to environmental stresses. To investigate its function in the drought stress response of wild jujube (Ziziphus jujuba Mill. var. spinosa), a typical eco-economic forest species, this study performed genome-wide identification and relevant analyses of R2R3-MYB genes. RESULTS: A total of 91 R2R3-MYB genes (designated as ZjMYB1 to ZjMYB91) were identified, which were unevenly distributed across 12 chromosomes. These genes mainly encode hydrophilic and unstable proteins, 97.8% of which are localized in the nucleus. Phylogenetic analysis classified these genes into 25 clades, showing evolutionary conservation and species-specific divergence with the R2R3-MYB protein family. The expansion of the ZjMYB family is mainly characterized by segmental duplication, and all duplicated gene pairs have undergone purifying selection. ZjMYBs are widely involved in plant growth and development as well as abiotic stress responses, with the highest expression level particularly in leaf tissues; a total of 13 genes were specifically annotated as water deficit response-related genes in drought stress and abscisic acid (ABA) signaling pathways. Integrating the above analyses together with transcriptome data and qRT-PCR validation results revealed that ZjMYB5, ZjMYB53, ZjMYB57 and ZjMYB85 function as core drought-responsive genes, which display both tissue-specific and time-dependent expression patterns under drought stress. CONCLUSIONS: This study systematically elucidated the functional characteristics and regulatory network of the R2R3-MYB gene family in wild jujube, providing critical genetic resources and a theoretical basis for dissecting the molecular mechanisms underlying drought tolerance in wild jujube and breeding drought-resistant cultivars.

Ziziphus↗

[Gene expression profiles on three kinds of genotype hepatitis C virus core protein in Huh-7 cell line with microarray analysis].

OBJECTIVE: To analyze three kinds of genotype hepatitis C virus (HCV) core protein expressed in human hepatoma (Huh-7) cell line and to recognize HCV core proteins biological function and its pathogenic mechanism. METHODS: The Huh-7 cell expressed three kinds of core proteins were established respectively. Affymetrix human gene chip was used for identifying the gene expression dependently on Affymetrix's protocol. All genes changed by 3 or 1.5 folds between the transfected cells and a control cells were further analyzed, and annotated by using NetAffx analysis through Affymetrix website and were categorized based on their biological processes. RESULTS: The HCV-1b core protein caused 16 genes up/down-regulated expression, of which the immune response genes of PF4V1 and SPP1 were up-regulated 3.4 or 4.4 folds respectively. The HCV-2a core protein had caused the immune response gene CXCL5 and apoptosis gene BTF a down-regulated expression of 3.4 and 3.1 folds respectively, but caused the apoptosis genes of HRK and LZTS1 an up-regulated expression of 3.2 and 3.4 folds respectively. As compared with HCV 1b or 2a core protein, HCV-4b core protein caused 111 genes expression changing and it had more obvious effects on gene expression. If we applied 1.5 fold change for a comparison gene expression, a few of the same gene expression profiles might be caused by these two core proteins. CONCLUSION: The three kinds of HCV core protein should have its own expression character and be mainly shown in immune responses, signal transduction, apoptosis, etc. It should be helpful for our recognizing the HCV core protein biological function and its pathogenic mechanism.

Carcinoma, Hepatocellular↗

The Candida Genome Database (CGD), a community resource for Candida albicans gene and protein information.

The Candida Genome Database (CGD) is a new database that contains genomic information about the opportunistic fungal pathogen Candida albicans. CGD is a public resource for the research community that is interested in the molecular biology of this fungus. CGD curators are in the process of combing the scientific literature to collect all C.albicans gene names and aliases; to assign gene ontology terms that describe the molecular function, biological process, and subcellular localization of each gene product; to annotate mutant phenotypes; and to summarize the function and biological context of each gene product in free-text description lines. CGD also provides community resources, including a reservation system for gene names and a colleague registry through which Candida researchers can share contact information and research interests. CGD is publicly funded (by NIH grant R01 DE15873-01 from the NIDCR) and is freely available at http://www.candidagenome.org/.

Candida albicans↗

Molecular cloning and functional expression of the first two specific insect myosuppressin receptors.

The Drosophila Genome Project database contains the sequences of two genes, CG8985 and CG13803, which are predicted to code for G protein-coupled receptors. We cloned the cDNAs corresponding to these genes and found that their gene structures had not been correctly annotated. We subsequently expressed the coding regions of the two corrected receptor genes in Chinese hamster ovary cells and found that each of them coded for a receptor that could be activated by low concentrations of Drosophila myosuppressin (EC50,4 x 10(-8) M). The insect myosuppressins are decapeptides that generally inhibit insect visceral muscles. Other tested Drosophila neuropeptides did not activate the two receptors. In addition to the two Drosophila myosuppressin receptors, we identified a sequence in the genomic database from the malaria mosquito Anopheles gambiae that also very likely codes for a myosuppressin receptor. To our knowledge, this paper is the first report on the molecular identification of specific insect myosuppressin receptors.

Amino Acid Sequence↗