Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,351 records · Page 75Linked to original sources

Confirmation of data mining based predictions of protein function.

MOTIVATION: A central problem in bioinformatics is the assignment of function to sequenced open reading frames (ORFs). The most common approach is based on inferred homology using a statistically based sequence similarity (SIM) method, e.g. PSI-BLAST. Alternative non-SIM based bioinformatic methods are becoming popular. One such method is Data Mining Prediction (DMP). This is based on combining evidence from amino-acid attributes, predicted structure and phylogenic patterns; and uses a combination of Inductive Logic Programming data mining, and decision trees to produce prediction rules for functional class. DMP predictions are more general than is possible using homology. In 2000/1, DMP was used to make public predictions of the function of 1309 Escherichia coli ORFs. Since then biological knowledge has advanced allowing us to test our predictions. RESULTS: We examined the updated (20.02.02) Riley group genome annotation, and examined the scientific literature for direct experimental derivations of ORF function. Both tests confirmed the DMP predictions. Accuracy varied between rules, and with the detail of prediction, but they were generally significantly better than random. For voting rules, accuracies of 75-100% were obtained. Twenty-one of these DMP predictions have been confirmed by direct experimentation. The DMP rules also have interesting biological explanations. DMP is, to the best of our knowledge, the first non-SIM based prediction method to have been tested directly on new data. AVAILABILITY: We have designed the "Genepredictions" database for protein functional predictions. This is intended to act as an open repository for predictions for any organism and can be accessed at http://www.genepredictions.org

Abstracting and Indexing↗

Bifunctional phosphoglucose/phosphomannose isomerase from the hyperthermophilic archaeon Pyrobaculum aerophilum.

ORF PAE1610 from the hyperthermophilic crenarchaeon Pyrobaculum aerophilum was first annotated as the conjectural pgi gene coding for hypothetical phosphoglucose isomerase (PGI). However, we have recently identified this ORF as the putative pgi/pmi gene coding for hypothetical bifunctional phosphoglucose/phosphomannose isomerase (PGI/PMI). To prove its coding function, ORF PAE1610 was overexpressed in Escherichia coli, and the recombinant enzyme was characterized. The 65-kDa homodimeric protein catalyzed the isomerization of both glucose-6-phosphate and mannose-6-phosphate to fructose-6-phosphate at similar catalytic rates, thus characterizing the enzyme as bifunctional PGI/PMI. The enzyme was extremely thermoactive; it had a temperature optimum for catalytic activity of about 100 degrees C and a melting temperature for thermal unfolding above 100 degrees C.

DNA, Archaeal↗

Biochemical properties and cellular localization of Plasmodium falciparum protein disulfide isomerase.

We have previously reported the isolation of a 52,000 M(r) protein (Pf52) displaying consensus sequences for thiol:disulfide oxidoreductases. Pf52 therefore represents the plasmodial protein disulfide isomerase (PDI). It has been renamed PfPDI and correlates to MAL8P1.17 in the annotated genome of P. falciparum (3D7 strain). Antibodies were raised against recombinant (His)(6)-tagged forms of PfPDI devoid of its signal peptide sequence, demonstrating a major co-localization of PfPDI with endoplasmic reticulum-resident proteins, PfBIP and PfERC, but not with the Golgi marker PfERD2. Recombinant PfPDI displayed typical biochemical functions of PDIs: oxidase/isomerase and reductase activities, as well as a chaperone-like behavior on the denaturated protein rhodanese. These activities were comparable to those measured for the purified native bovine PDI and the human recombinant PDI. The antiplasmodial compound DS61 does inhibit the recombinant PfPDI oxidase/isomerase activity but not that of the human recombinant PDI, suggesting structural differences between both enzymes. However, a discrepancy between the inhibitory activity of DS61 on the recombinant PfPDI (IC(50) of 430 microM) and its in vitro antiplasmodial activity (IC(50) of 0.1 microM) was observed, suggesting that PfPDI is not the only target of DS61. Taking into account its biochemical properties and its intracellular localization, the involvement of PfPDI in the parasite protein folding is discussed, as well as its potential for the development of alternative antimalarial chemotherapy strategies.

Animals↗

The genomic view of genes responsive to the antagonistic phytohormones, abscisic acid, and gibberellin.

We now have the various genomics tools for monocot (Oryza sativa) and a dicot (Arabidopsis thaliana) plant. Plant is not only a very important agricultural resource but also a model organism for biological research. It is important that the interaction between ABA and GA is investigated for controlling the transition from embryogenesis to germination in seeds using genomics tools. These studies have investigated the relationship between dormancy and germination using genomics tools. Genomics tools identified genes that had never before been annotated as ABA- or GA-responsive genes in plant, detected new interactions between genes responsive to the two hormones, comprehensively characterized cis-elements of hormone-responsive genes, and characterized cis-elements of rice and Arabidopsis. In these research, ABA- and GA-regulated genes have been classified as functional proteins (proteins that probably function in stress or PR tolerance) and regulatory proteins (protein factors involved in further regulation of signal transduction). Comparison between ABA and/or GA-responsive genes in rice and those in Arabidopsis has shown that the cis-element has specificity in each species. cis-Elements for the dehydration-stress response have been specified in Arabidopsis but not in rice. cis-Elements for protein storage are remarkably richer in the upstream regions of the rice gene than in those of Arabidopsis.

Abscisic Acid↗

Structural and functional investigation of a putative archaeal selenocysteine synthase.

Bacterial selenocysteine synthase converts seryl-tRNA(Sec) to selenocysteinyl-tRNA(Sec) for selenoprotein biosynthesis. The identity of this enzyme in archaea and eukaryotes is unknown. On the basis of sequence similarity, a conserved open reading frame has been annotated as a selenocysteine synthase gene in archaeal genomes. We have determined the crystal structure of the corresponding protein from Methanococcus jannaschii, MJ0158. The protein was found to be dimeric with a distinctive domain arrangement and an exposed active site, built from residues of the large domain of one protomer alone. The shape of the dimer is reminiscent of a substructure of the decameric Escherichia coli selenocysteine synthase seen in electron microscopic projections. However, biochemical analyses demonstrated that MJ0158 lacked affinity for E. coli seryl-tRNA(Sec) or M. jannaschii seryl-tRNA(Sec), and neither substrate was directly converted to selenocysteinyl-tRNA(Sec) by MJ0158 when supplied with selenophosphate. We then tested a hypothetical M. jannaschii O-phosphoseryl-tRNA(Sec) kinase and demonstrated that the enzyme converts seryl-tRNA(Sec) to O-phosphoseryl-tRNA(Sec) that could constitute an activated intermediate for selenocysteinyl-tRNA(Sec) production. MJ0158 also failed to convert O-phosphoseryl-tRNA(Sec) to selenocysteinyl-tRNA(Sec). In contrast, both archaeal and bacterial seryl-tRNA synthetases were able to charge both archaeal and bacterial tRNA(Sec) with serine, and E. coli selenocysteine synthase converted both types of seryl-tRNA(Sec) to selenocysteinyl-tRNA(Sec). These findings demonstrate that a number of factors from the selenoprotein biosynthesis machineries are cross-reactive between the bacterial and the archaeal systems but that MJ0158 either does not encode a selenocysteine synthase or requires additional factors for activity.

Amino Acid Sequence↗

A chromosome-level genome assembly and annotation of Cercis chuniana (Fabaceae).

The genus Cercis L., at the base of the subfamily Cercidoideae of Fabaceae, is known for its ecological adaptability and significant medicinal, ornamental, and economic value. However, the lack of a high-quality genome hinders the understanding of the evolution of Cercis and Fabaceae. In this study, we present a chromosome-level genome of Cercis chuniana by combining Illumina short reads, PacBio HiFi long reads, and Hi-C data. The final genome size is 355.53 Mb, consisting of 12 contigs with a N50 of 42.34 Mb. Notably, 344.24 Mb, corresponding to 96.82% of the genome, was anchored to seven chromosomes. The assembly comprises 24.83% repetitive sequences, including 19.32% long terminal repeats. Additionally, a total of 33,837 protein-coding genes were predicted in the genome, with 32,709 (96.67%) genes successfully annotated. The high-quality genome assembly of C. chuniana not only bridges the existing gap in genomic data and offers important resources for molecular studies of this species, but also provides essential insights for future studies on speciation, functional and comparative genomics within the Fabaceae family.

Genome, Plant↗

Genomic anatomy of the Tyrp1 (brown) deletion complex.

Chromosome deletions in the mouse have proven invaluable in the dissection of gene function. The brown deletion complex comprises >28 independent genome rearrangements, which have been used to identify several functional loci on chromosome 4 required for normal embryonic and postnatal development. We have constructed a 172-bacterial artificial chromosome contig that spans this 22-megabase (Mb) interval and have produced a contiguous, finished, and manually annotated sequence from these clones. The deletion complex is strikingly gene-poor, containing only 52 protein-coding genes (of which only 39 are supported by human homologues) and has several further notable genomic features, including several segments of >1 Mb, apparently devoid of a coding sequence. We have used sequence polymorphisms to finely map the deletion breakpoints and identify strong candidate genes for the known phenotypes that map to this region, including three lethal loci (l4Rn1, l4Rn2, and l4Rn3) and the fitness mutant brown-associated fitness (baf). We have also characterized misexpression of the basonuclin homologue, Bnc2, associated with the inversion-mediated coat color mutant white-based brown (B(w)). This study provides a molecular insight into the basis of several characterized mouse mutants, which will allow further dissection of this region by targeted or chemical mutagenesis.

Animals↗

Comparative genomics and experimental characterization of N-acetylglucosamine utilization pathway of Shewanella oneidensis.

We used a comparative genomics approach implemented in the SEED annotation environment to reconstruct the chitin and GlcNAc utilization subsystem and regulatory network in most proteobacteria, including 11 species of Shewanella with completely sequenced genomes. Comparative analysis of candidate regulatory sites allowed us to characterize three different GlcNAc-specific regulons, NagC, NagR, and NagQ, in various proteobacteria and to tentatively assign a number of novel genes with specific functional roles, in particular new GlcNAc-related transport systems, to this subsystem. Genes SO3506 and SO3507, originally annotated as hypothetical in Shewanella oneidensis MR-1, were suggested to encode novel variants of GlcN-6-P deaminase and GlcNAc kinase, respectively. Reconstitution of the GlcNAc catabolic pathway in vitro using these purified recombinant proteins and GlcNAc-6-P deacetylase (SO3505) validated the entire pathway. Kinetic characterization of GlcN-6-P deaminase demonstrated that it is the subject of allosteric activation by GlcNAc-6-P. Consistent with genomic data, all tested Shewanella strains except S. frigidimarina, which lacked representative genes for the GlcNAc metabolism, were capable of utilizing GlcNAc as the sole source of carbon and energy. This study expands the range of carbon substrates utilized by Shewanella spp., unambiguously identifies several genes involved in chitin metabolism, and describes a novel variant of the classical three-step biochemical conversion of GlcNAc to fructose 6-phosphate first described in Escherichia coli.

Acetylglucosamine↗

Isolation and transcription profiling of low-O2 stress-associated cDNA clones from the flooding-stress-tolerant FR13A rice genotype.

BACKGROUND: and Aims Flooding stress leads to a significant reduction in transcription and translation of genes involved in basal metabolism of plants. However, specific genes are noted to be up-regulated in this response. With the aim of isolating genes that might be specifically involved in flooding stress-tolerance mechanism(s), two subtractive cDNA libraries for the flooding-stress-tolerant rice genotype FR13A have been constructed, namely the single and double subtraction libraries (SSL and DSL, respectively). METHODS: To construct the SSL, mRNAs present in the unstressed control FR13A roots were subtracted from the mRNA pool present in low O2-stressed roots of FR13A rice seedlings. The DSL was constructed from mRNAs isolated from the roots of low O2-stressed FR13A rice seedlings from which pools of low-O2-stress up-regulated mRNAs from Pusa Basmati 1 and constitutively expressed mRNAs from FR13A roots were subtracted. RESULTS: In all, 400 and 606 cDNA clones were obtained from the SSL and DSL, respectively. Global transcript profiling by reverse northern analysis revealed that a large number of clones from these libraries were up-regulated by anaerobic stress. Importantly, selective up-regulated clones showed characteristic cultivar- and tissue-specific expression profiles. Sequencing and annotation of the up-regulated clones revealed that specific signal proteins, hexose transporters, ion channel transporters, RNA-binding proteins and transcription factor proteins possibly play important roles in the response of rice to flooding stress. Also a significant number of novel cDNA clones was noted in these libraries. CONCLUSIONS: It appears that cellular functions such as signalling, sugar and ion transport and transcript stability play an important role in conferring higher flooding tolerance in the FR13A rice type.

DNA, Complementary↗

FGDB: a comprehensive fungal genome resource on the plant pathogen Fusarium graminearum.

The MIPS Fusarium graminearum Genome Database (FGDB) is a comprehensive genome database on one of the most devastating fungal plant pathogens of wheat and barley. FGDB provides information on two gene sets independently derived by automated annotation of the F.graminearum genome sequence. A complete manually revised gene set will be completed within the near future. The initial results of systematic manual correction of gene calls are already part of the current gene set. The database can be accessed to retrieve information from bioinformatics analyses and functional classifications of the proteins. The data are also organized in the well established MIPS catalogs and novel query techniques are available to search the data. The comprehensive set of gene calls was also used for the design of an Affymetrix GeneChip. The resource is accessible on http://mips.gsf.de/genre/proj/fusarium/.

Databases, Genetic↗

Identification of mammalian microRNA host genes and transcription units.

To derive a global perspective on the transcription of microRNAs (miRNAs) in mammals, we annotated the genomic position and context of this class of noncoding RNAs (ncRNAs) in the human and mouse genomes. Of the 232 known mammalian miRNAs, we found that 161 overlap with 123 defined transcription units (TUs). We identified miRNAs within introns of 90 protein-coding genes with a broad spectrum of molecular functions, and in both introns and exons of 66 mRNA-like noncoding RNAs (mlncRNAs). In addition, novel families of miRNAs based on host gene identity were identified. The transcription patterns of all miRNA host genes were curated from a variety of sources illustrating spatial, temporal, and physiological regulation of miRNA expression. These findings strongly suggest that miRNAs are transcribed in parallel with their host transcripts, and that the two different transcription classes of miRNAs ('exonic' and 'intronic') identified here may require slightly different mechanisms of biogenesis.

Base Sequence↗

Peripheral genotype-phenotype correlations in Asian Indians with type 2 diabetes mellitus.

OBJECTIVE: A genome-wide scan of gene expression in leucocytes in Asian Indians with type 2 diabetes was performed and correlated with their known phenotype. METHODS: Microarray gene profiling of 13,474 sequence-verified, non-redundant human cDNAs was done to study leukocyte gene expression in Asian Indians with type 2 diabetes (DM: n=3) and matched controls (n=3). RESULTS: Significant differential expression (fold change <0.3 or >3) was noted for 897 genes in DM vs. controls. The 147 known genes in this category belonged to following broad functional groups (%): enzyme (32), nucleic acid binding (22), ligand binding or carrier (10), signal transducer (9), transporter (7), structural protein (6), cell adhesion (3), tumor suppressor (3), transcription factor binding (2), enzyme inhibitor (2), chaperone (2), cell cycle regulator (1), and defense/immunity protein (1). The 20 genes with at least a 3-fold change, annotated with known phenotypic associations in the current gene databank (phenotype association, fold change) were aspartoacylase (Canavan disease, 9.96), growth hormone receptor (Laron dwarfism, idiopathic short stature, 8.25), lipoprotein lipase (familial chylomicronemia syndrome, lipoprotein lipase deficiency, 8.00), vitamin D (1,25- dihydroxyvitamin D3) receptor (involutional osteoporosis, vitamin D resistant rickets, 7.94), intercellular adhesion molecule 1 human rhinovirus receptor (cerebral malaria susceptibility, 7.16), peroxisomal membrane protein 3 35-kDa (Refsum disease, infantile form, Zellweger syndrome-3, 6.00), Bardet-Biedl syndrome 2 (Bardet-Biedl syndrome, 5.87), ribosomal protein S19 (Diamond Blackfan anemia, 5.85), apolipoprotein C-III (hypertriglyceridemia, 5.44), argininosuccinate lyase (argininosuccinicaciduria, 5.22), myosin VA (Griscelli syndrome-type pigmentary dilution with mental retardation, 4.92), lysozyme (renal amyloidosis, 4.17), SAM domain, SH3 domain and nuclear localisation signals 1 (Cherubism, 4.12 ), von Hippel-Lindau syndrome (hemangioblastoma, cerebellar, somatic, von Hippel-Lindau syndrome, 3.94), early-onset breast cancer 1 (BRCA1, papillary serous carcinoma of the peritoneum, 3.73), UDP-N-acetylglucosamine-2-epimerase/N-acetylmannosamine kinase (inclusion body myopathy, autosomal recessive, sialuria, 3.53), apolipoprotein A-I (amyloidosis, 3 or more types, hypoalphalipoproteinemia, 3.29), midline 1 Opitz/BBB syndrome (Opitz G syndrome, type I, 3.28), ATPase, Na+/K+ transporting, alpha 2 (+) polypeptide (familial hemiplegic migraine, 3.05). Canavan disease, Zellweger syndrome, infantile Refsum disease, Griscelli syndrome, cherubism, breast cancer, peritoneal papillary serous carcinoma, Opitz G/BBB syndrome, and familial hemiplegic migraine (FHM) are phenotypes not previously reported in association with type 2 DM, but whose underlying genes were up-regulated in this peripheral genome scan of Asian Indians. CONCLUSION: Rare and/or previously unknown phenotypes linked to known genes with significant differential expression in type 2 DM are reported. Further testing of heterogeneity in diabetes phenotype syndromes may reveal common pathogenic mechanisms and potential candidate genes responsible for type 2 DM.

Asian People↗

G protein-coupled receptor genes in the FANTOM2 database.

G protein-coupled receptors (GPCRs) comprise the largest family of receptor proteins in mammals and play important roles in many physiological and pathological processes. Gene expression of GPCRs is temporally and spatially regulated, and many splicing variants are also described. In many instances, different expression profiles of GPCR gene are accountable for the changes of its biological function. Therefore, it is intriguing to assess the complexity of the transcriptome of GPCRs in various mammalian organs. In this study, we took advantage of the FANTOM2 (Functional Annotation Meeting of Mouse cDNA 2) project, which aimed to collect full-length cDNAs inclusively from mouse tissues, and found 410 candidate GPCR cDNAs. Clustering of these clones into transcriptional units (TUs) reduced this number to 213. Out of these, 165 genes were represented within the known 308 GPCRs in the Mouse Genome Informatics (MGI) resource. The remaining 48 genes were new to mouse, and 14 of them had no clear mammalian ortholog. To dissect the detailed characteristics of each transcript, tissue distribution pattern and alternative splicing were also ascertained. We found many splicing variants of GPCRs that may have a relevance to disease occurrence. In addition, the difficulty in cloning tissue-specific and infrequently transcribed GPCRs is discussed further.

Alternative Splicing↗

Discovery of 342 putative new genes from the analysis of 5'-end-sequenced full-length-enriched cDNA human transcripts.

In this work we describe the process that, starting with the production of human full-length-enriched cDNA libraries using the CAP-Trapper method, led us to the discovery of 342 putative new human genes. Twenty-three thousand full-length-enriched clones, obtained from various cell lines and tissues in different developmental stages, were 5'-end sequenced, allowing the identification of a pool of 5300 unique cDNAs. By comparing these sequences to various human and vertebrate nucleotide databases we found that about 40% of our clones extended previously annotated 5' ends, 662 clones were likely to represent splice variants of known genes, and finally 342 clones remained unknown, with no or poor functional annotation. cDNA-microarray gene expression analysis showed that 260 of 342 unknown clones are expressed in at least one cell line and/or tissue. Further analysis of their sequences and the corresponding genomic locations allowed us to conclude that most of them represent potential novel genes, with only a small fraction having protein-coding potential.

5' Flanking Region↗

Integrated analysis of protein composition, tissue diversity, and gene regulation in mouse mitochondria.

Mitochondria are tailored to meet the metabolic and signaling needs of each cell. To explore its molecular composition, we performed a proteomic survey of mitochondria from mouse brain, heart, kidney, and liver and combined the results with existing gene annotations to produce a list of 591 mitochondrial proteins, including 163 proteins not previously associated with this organelle. The protein expression data were largely concordant with large-scale surveys of RNA abundance and both measures indicate tissue-specific differences in organelle composition. RNA expression profiles across tissues revealed networks of mitochondrial genes that share functional and regulatory mechanisms. We also determined a larger "neighborhood" of genes whose expression is closely correlated to the mitochondrial genes. The combined analysis identifies specific genes of biological interest, such as candidates for mtDNA repair enzymes, offers new insights into the biogenesis and ancestry of mammalian mitochondria, and provides a framework for understanding the organelle's contribution to human disease.

Animals↗

The ABC of ABCS: a phylogenetic and functional classification of ABC systems in living organisms.

ATP binding cassette (ABC) systems constitute one of the most abundant superfamilies of proteins. They are involved not only in the transport of a wide variety of substances, but also in many cellular processes and in their regulation. In this paper, we made a comparative analysis of the properties of ABC systems and we provide a phylogenetic and functional classification. This analysis will be helpful to accurately annotate ABC systems discovered during the sequencing of the genome of living organisms and to identify the partners of the ABC ATPases.

ATP-Binding Cassette Transporters↗

ORFDB: an information resource linking scientific content to a high-quality Open Reading Frame (ORF) collection.

The ORFDB (http://orf.invitrogen.com/) represents an ongoing effort at Invitrogen Corporation to integrate relevant scientific data with an evolving collection of human and mouse Open Reading Frame (ORF) clones (Ultimate ORF Clones). The ORFDB serves as a central data warehouse enabling researchers to search the ORF collection through its web portal ORFBrowser, allowing researchers to find the Ultimate ORF clones by blast, keyword, GenBank accession, gene symbol, clone ID, Unigene ID, LocusLink ID or through functional relationships by browsing the collection via the Gene Ontology (GO) Browser. As of October 2003, the ORFDB contains 6200 human and 2870 mouse Ultimate ORF clones. All Ultimate ORF clones have been fully sequenced with high quality, and are matched to public reference protein sequences. In addition, the cloned ORFs have been extensively annotated across six categories: Gene, ORF, Clone Format, Protein, SNP and Genomic links, with the information assembled in a format termed the ORFCard. The ORFCard represents an information repository that documents the sequence quality, alignment with respect to public protein sequences, and the latest publicly available information associated with each human and mouse gene represented in the collection.

Animals↗

The REFOLD database: a tool for the optimization of protein expression and refolding.

A large proportion of proteins expressed in Escherichia coli form inclusion bodies and thus require renaturation to attain a functional conformation for analysis. In this process, identifying and optimizing the refolding conditions and methodology is often rate limiting. In order to address this problem, we have developed REFOLD, a web-accessible relational database containing the published methods employed in the refolding of recombinant proteins. Currently, REFOLD contains >300 entries, which are heavily annotated such that the database can be searched via multiple parameters. We anticipate that REFOLD will continue to grow and eventually become a powerful tool for the optimization of protein renaturation. REFOLD is freely available at http://refold.med.monash.edu.au.

Databases, Protein↗