Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Chromosome-level genome assembly of Sinocyclocheilus jii based on PacBio HiFi and Hi-C sequencing.

Sinocyclocheilus jii, a cavefish species endemic to China, belongs to the genus Sinocyclocheilus within the family Cyprinidae. Species within this genus exhibit significant morphological differentiation, making it not only the most species-rich genus within Cyprinidae in China but also the most diverse group of cavefishes worldwide. However, the limited availability of genomic resources has limited investigations into the genetic basis of trait variations, phylogenetic relationships, and adaptive evolution in this genus. In this study, we assembled a chromosome-level reference genome for S. jii by integrating PacBio HiFi long reads, Illumina short reads, and Hi-C sequencing data. Flow cytometry was used to estimate the genome size prior to assembly, providing a key step in technical validation. The final genome assembly spans 1.75 Gb with a contig N50 of 35.0 Mb. Using Hi-C sequencing data, the assembled scaffolds were successfully anchored to 50 chromosomes. The completeness of the chromosome-level assembly was estimated at 98.9% by BUSCO analysis. Genome annotation identified 855.5 Mb of repetitive sequences and predicted a total of 52,867 protein-coding genes, of which 51,932 genes were functionally annotated. This study presents a high-quality chromosome-level genome assembly and annotation of S. jii, providing a fundamental genomic resource for future phylogenetic and evolutionary studies.

Animals↗

Information assessment on predicting protein-protein interactions.

BACKGROUND: Identifying protein-protein interactions is fundamental for understanding the molecular machinery of the cell. Proteome-wide studies of protein-protein interactions are of significant value, but the high-throughput experimental technologies suffer from high rates of both false positive and false negative predictions. In addition to high-throughput experimental data, many diverse types of genomic data can help predict protein-protein interactions, such as mRNA expression, localization, essentiality, and functional annotation. Evaluations of the information contributions from different evidences help to establish more parsimonious models with comparable or better prediction accuracy, and to obtain biological insights of the relationships between protein-protein interactions and other genomic information. RESULTS: Our assessment is based on the genomic features used in a Bayesian network approach to predict protein-protein interactions genome-wide in yeast. In the special case, when one does not have any missing information about any of the features, our analysis shows that there is a larger information contribution from the functional-classification than from expression correlations or essentiality. We also show that in this case alternative models, such as logistic regression and random forest, may be more effective than Bayesian networks for predicting interactions. CONCLUSIONS: In the restricted problem posed by the complete-information subset, we identified that the MIPS and Gene Ontology (GO) functional similarity datasets as the dominating information contributors for predicting the protein-protein interactions under the framework proposed by Jansen et al. Random forests based on the MIPS and GO information alone can give highly accurate classifications. In this particular subset of complete information, adding other genomic data does little for improving predictions. We also found that the data discretizations used in the Bayesian methods decreased classification performance.

Artificial Intelligence↗

Proteome of Methanosarcina acetivorans Part II: comparison of protein levels in acetate- and methanol-grown cells.

Methanosarcina acetivorans is an archaeon isolated from marine sediments which utilizes a diversity of substrates for growth and methanogenesis. Part I of a two-part investigation has profiled proteins of this microorganism cultured with both methanol and acetate as growth substrates, utilizing two-dimensional gel electrophoresis and MALDI-TOF-TOF mass spectrometry. In this report, Part II, the analyses were extended to identify 34 proteins found to be present in different amounts between methanol- and acetate-grown M. acetivorans. Among these proteins are enzymes which function in pathways for methanogenesis from either acetate or methanol. Several of the 34 proteins were determined to have redundant functions based on annotations of the genomic sequence. Enzymes which function in ATP synthesis and steps common to both methanogenic pathways were elevated in acetate- versus methanol-grown cells, whereas enzymes that have a more general function in protein synthesis were in greater amounts in methanol- compared to acetate-grown cells. Several group I chaperonins were present in greater amounts in methanol- versus acetate-grown cells, whereas lower amounts of several stress related proteins were found in methanol- versus acetate-grown cells. The potential physiological basis for these novel patterns of protein synthesis are discussed.

Acetates↗

The tegument surface membranes of the human blood parasite Schistosoma mansoni: a proteomic analysis after differential extraction.

The blood fluke Schistosoma mansoni can live for years in the hepatic portal system of its human host and so must possess very effective mechanisms of immune evasion. The key to understanding how these operate lies in defining the molecular organisation of the exposed parasite surface. The adult worm is covered by a syncytial tegument, bounded externally by a plasma membrane and overlain by a laminate secretion, the membranocalyx. In order to determine the protein composition of this surface, the membranes were detached using a freeze/thaw technique and enriched by sucrose density gradient centrifugation. The resulting preparation was sequentially extracted with three reagents of increasing solubilising power. The extracts were separated by 2-DE and their protein constituents were identified by MS/MS, yielding predominantly cytosolic, cytoskeletal and membrane-associated proteins, respectively. After extraction, the final pellet containing membrane-spanning proteins was processed by liquid chromatographic techniques before MS. Transporters for sugars, amino acids, ions and other solutes were found together with membrane enzymes and proteins concerned with membrane structure. The proteins identified were categorised by their function and putative location on the basis of their homology with annotated proteins in other organisms.

Animals↗

The characterisation of novel secreted Ly-6 proteins from rat urine by the combined use of two-dimensional gel electrophoresis, microbore high performance liquid chromatography and expressed sequence tag data.

A proteomic study of rat urine was undertaken using two-dimensional gel electrophoresis, microbore high performance liquid chromatography, mass spectrometry and N-terminal sequencing. Five known urinary proteins were identified but two novel peptide fragments matched a large number of rat expressed sequence tags (ESTs) from a liver library. By combining protein chemical and nucleotide data, two 101-residue open reading frames with 90% amino acid identity were determined, rat urinary protein 1 (RUP-1) and RUP-2. The data established signal peptide removal and provided evidence for N-glycosylation. A third related sequence, rat spleen protein (RSP-1) was confirmed from EST searches. These three proteins have been submitted to SWISS-PROT as P81827, P81828 and Q9QXN2, respectively. A fourth novel homologue was found in porcine and bovine ESTs from embryo libraries. Alignment with known homologues showed conserved cysteine positions characteristic of a secreted subfamily of Ly-6 proteins. In two cases, antineoplastic urinary protein and caltrin, these homologues have unverified functional annotations. The RUP sequences showed high scoring matches to three unrelated rat mRNAs subsequently established to be chimeric. Two of these share extended sectional identity to RUP-1 but the third may represent another novel Ly-6 homologue. These chimeras have caused serious annotation errors in secondary databases.

Amino Acid Sequence↗

Similarity networks of protein binding sites.

An increasing attention has been dedicated to the characterization of complex networks within the protein world. This work is reporting how we uncovered networked structures that reflected the structural similarities among protein binding sites. First, a 211 binding sites dataset has been compiled by removing the redundant proteins in the Protein Ligand Database (PLD) (http://www-mitchell.ch.cam.ac.uk/pld/). Using a clique detection algorithm we have performed all-against-all binding site comparisons among the 211 available ones. Within the set of nodes representing each binding site an edge was added whenever a pair of binding sites had a similarity higher than a threshold value. The generated similarity networks revealed that many nodes had few links and only few were highly connected, but due to the limited data available it was not possible to definitively prove a scale-free architecture. Within the same dataset, the binding site similarity networks were compared with the networks of sequence and fold similarity networks. In the protein world, indications were found that structure is better conserved than sequence, but on its own, sequence was better conserved than the subset of functional residues forming the binding site. Because a binding site is strongly linked with protein function, the identification of protein binding site similarity networks could accelerate the functional annotation of newly identified genes. In view of this we have discussed several potential applications of binding site similarity networks, such as the construction of novel binding site classification databases, as well as the implications for protein molecular design in general and computational chemogenomics in particular.

Animals↗

The complete genome sequence of Escherichia coli K-12.

The 4,639,221-base pair sequence of Escherichia coli K-12 is presented. Of 4288 protein-coding genes annotated, 38 percent have no attributed function. Comparison with five other sequenced microbes reveals ubiquitous as well as narrowly distributed gene families; many families of similar genes within E. coli are also evident. The largest family of paralogous proteins contains 80 ABC transporters. The genome as a whole is strikingly organized with respect to the local direction of replication; guanines, oligonucleotides possibly related to replication and recombination, and most genes are so oriented. The genome also contains insertion sequence (IS) elements, phage remnants, and many other patches of unusual composition indicating genome plasticity through horizontal transfer.

Bacterial Proteins↗

Mutational data integration in gene-oriented files of the Hermansky-Pudlak Syndrome database.

Hermansky-Pudlak Syndrome (HPS) is a genetically heterogeneous disorder characterized by oculocutaneous albinism and prolonged bleeding due to abnormal vesicle trafficking to lysosomes and related organelles such as melanosomes and platelet dense granules. This HPS database (HPSD; http://liweilab.genetics.ac.cn/HPSD/) provides integrated, annotatory, and curative data that is distributed in a variety of public databases or predicted by bioinformatics servers for the recently cloned human and mouse HPS genes, as well as for the genes responsible for HPSrelated syndromes, such as ChediakHigashi Syndrome (CHS), Griscelli syndrome (GS), oculocutaneous albinism (OCA), Usher syndrome type 1B (USH1B), and ocular albinism (OA). The HPSD is designed by using a unique GeneOriented File (GOF) format. Seven blocks (genomic, transcript, protein, function, mutation, phenotype, and reference) are carefully annotated in each userfriendly GOF entry. The HPSD emphasizes paired human and mouse GOF entries. The genes included in this database (currently 58 in total) are arbitrarily divided into four categories: 1) Human and Mouse HPS, 2) Mouse HPS Only, 3) Putative Mouse or Human HPS, and 4) HPS Related Syndromes. All the mutations in these genes are integrated in the GOFs. We expect that these very informative and peerreviewed GOFs will be shortcuts to utilize the webbased information for the emerging interdisciplinary studies of HPS.

Animals↗

SubtiList: the reference database for the Bacillus subtilis genome.

SubtiList is the reference database dedicated to the genome of Bacillus subtilis 168, the paradigm of Gram-positive endospore-forming bacteria. Developed in the framework of the B.subtilis genome project, SubtiList provides a curated dataset of DNA and protein sequences, combined with the relevant annotations and functional assignments. Information about gene functions and products is continuously updated by linking relevant bibliographic references. Recently, sequence corrections arising from both systematic verifications and submissions by individual scientists were included in the reference genome sequence. SubtiList is based on a generic relational data schema and a World Wide Web interface developed for the handling of bacterial genomes, called GenoList. The World Wide Web interface was designed to allow users to easily browse through genome data and retrieve information according to common biological queries. SubtiList also provides more elaborate tools, such as pattern searching, which are tightly connected to the overall browsing system. SubtiList is accessible at http://genolist.pasteur.fr/SubtiList/. Similar bacterial databases are accessible at http://genolist.pasteur.fr/.

Bacillus subtilis↗

The SBASE protein domain library, release 6.0: a collection of annotated protein sequence segments.

The sixth release of the SBASE protein domain library sequences contains 130 703 annotated and crossreferenced entries corresponding to structural, functional, ligand-binding and topogenic segments of proteins. The entries were grouped based on standard names (2312 groups) and futher classified on the basis of the BLAST similarity (2463 clusters). Automated searching with BLAST and a new sequence-plot representation of local domain similarities are available at the WWW-server http://www.icgeb.trieste.it/sbase. A mirror site is at http://sbase.abc.hu/sbase. The database is freely available by anonymous 'ftp' file transfer from ftp.icgeb.trieste.it

Amino Acid Sequence↗

The identification of nucleic acid-interacting proteins using a simple proteomics-based approach that directly incorporates the electrophoretic mobility shift assay.

Proteins that interact with nucleic acids are central to numerous cellular processes, and their continuing characterization represents one of the foremost challenges in the postgenomic era. Here we describe a simple proteomics-based approach for the identification by mass spectrometry of proteins in crude extracts that interact with nucleic acids. It incorporates the electrophoretic mobility shift assay and is based on the finding that when a protein forms a complex with nucleic acid its electrophoretic mobility is affected as well as that of the nucleic acid. Our method should greatly reduce and in some cases may even eliminate the need for extensive protein purification and as such should contribute significantly to the functional annotation of the proteome. Furthermore it requires no prior knowledge of the molecular mass, quaternary structure, or pI of the interacting protein. Proof of principle is demonstrated using a recently discovered transcription factor; however, the approach should also have application in the identification of proteins that interact with RNA.

Amino Acid Sequence↗

Protein Information Resource: a community resource for expert annotation of protein data.

The Protein Information Resource, in collaboration with the Munich Information Center for Protein Sequences (MIPS) and the Japan International Protein Information Database (JIPID), produces the most comprehensive and expertly annotated protein sequence database in the public domain, the PIR-International Protein Sequence Database. To provide timely and high quality annotation and promote database interoperability, the PIR-International employs rule-based and classification-driven procedures based on controlled vocabulary and standard nomenclature and includes status tags to distinguish experimentally determined from predicted protein features. The database contains about 200,000 non-redundant protein sequences, which are classified into families and superfamilies and their domains and motifs identified. Entries are extensively cross-referenced to other sequence, classification, genome, structure and activity databases. The PIR web site features search engines that use sequence similarity and database annotation to facilitate the analysis and functional identification of proteins. The PIR-Inter-national databases and search tools are accessible on the PIR web site at http://pir.georgetown.edu/ and at the MIPS web site at http://www.mips.biochem.mpg.de. The PIR-International Protein Sequence Database and other files are also available by FTP.

Computational Biology↗

FunSpec: a web-based cluster interpreter for yeast.

BACKGROUND: For effective exposition of biological information, especially with regard to analysis of large-scale data types, researchers need immediate access to multiple categorical knowledge bases and need summary information presented to them on collections of genes, as opposed to the typical one gene at a time. RESULTS: We present here a web-based tool (FunSpec) for statistical evaluation of groups of genes and proteins (e.g. co-regulated genes, protein complexes, genetic interactors) with respect to existing annotations (e.g. functional roles, biochemical properties, localization). FunSpec is available online at http://funspec.med.utoronto.ca CONCLUSION: FunSpec is helpful for interpretation of any data type that generates groups of related genes and proteins, such as gene expression clustering and protein complexes, and is useful for predictive methods employing "guilt-by-association."

Cluster Analysis↗

The TIGR Plant Transcript Assemblies database.

The TIGR Plant Transcript Assemblies (TA) database (http://plantta.tigr.org) uses expressed sequences collected from the NCBI GenBank Nucleotide database for the construction of transcript assemblies. The sequences collected include expressed sequence tags (ESTs) and full-length and partial cDNAs, but exclude computationally predicted gene sequences. The TA database includes all plant species for which more than 1000 EST or cDNA sequences are publicly available. The EST and cDNA sequences are first clustered based on an all-versus-all pairwise sequence comparison, followed by the generation of consensus sequences (TAs) from individual clusters. The clustering and assembly procedures use the TGICL tool, Megablast and the CAP3 assembler. The UniProt Reference Clusters (UniRef100) protein database is used as the reference database for the functional annotation of the assemblies. The transcription orientation of each TA is determined based on the orientation of the alignment with the best protein hit. The TA sequences and annotation are available via web interfaces and FTP downloads. Assemblies can be retrieved by a text-based keyword search or a sequence-based BLAST search. The current version of the TA database is Release 2 (July 17, 2006) and includes a total of 215 plant species.

DNA, Complementary↗

Two-dimensional gel electrophoresis maps of the proteome and phosphoproteome of primitively cultured rat mesangial cells.

Mesangial cells (MC) play an important role in maintaining the structure and function of the glomerulus. The proliferation of MC is a prominent feature of many kinds of glomerular disease. The first reference 2-DE maps of rat mesangial cells (RMC), stained with silver staining or Pro-Q Diamond dye, have been established here to describe the proteome and phosphoproteome of RMC, respectively. A total of 157 selected protein spots, corresponding to 118 unique proteins, have been identified by MALDI-TOF-MS or LC-ESI-IT-MS/MS, in which 37 protein spots representing 28 unique proteins have also been stained with Pro-Q Diamond, indicating that they are in phosphorylated forms. All the identified proteins were bioinformatically annotated in detail according to their physiochemical characteristics, subcellular location, and function. Most of the separated or identified protein spots are distributed in the area of mass 10-70 kDa and pI 5.0-8.0. The identified proteins include mainly cytoplasmic and nuclear proteins and some mitochondrial, endoplasmic reticulum, and membrane proteins. These proteins are classified into different functional groups such as structure and mobility proteins (21.2%), metabolic enzymes (16.9%), protein folding and metabolism proteins (13.6%), signaling proteins (14.4%), heat-shock proteins (7.6%), and other functional proteins (12.7%). While structure and mobility proteins are mostly represented by protein spots with high abundance, signaling proteins are mostly represented by protein spots with relatively low abundance. Such a 2-DE database for RMC, especially with many signaling proteins and phosphoproteins characterized, will provide a valuable resource for comparative proteomics analysis of normal and pathologic conditions affecting MC function or pathologic progress.

Animals↗

Reannotation of Shewanella oneidensis genome.

As more and more complete bacterial genome sequences become available, the genome annotation of previously sequenced genomes may become quickly outdated. This is primarily due to the discovery and functional characterization of new genes. We have reannotated the recently published genome of Shewanella oneidensis with the following results: 51 new genes have been identified, and functional annotation has been added to the 97 genes, including 15 new and 82 existing ones with previously unassigned function. The identification of new genes was achieved by predicting the protein coding regions using the HMM-based program GeneMark.hmm. Subsequent comparison of the predicted gene products to the non-redundant protein database using BLAST and the COG (Clusters of Orthologous Groups) database using COGNITOR provided for the functional annotation.

Algorithms↗

Augur--a computational pipeline for whole genome microbial surface protein prediction and classification.

UNLABELLED: The analysis of protein function is a challenge and a major bottleneck towards well-annotated and analysed microbial genomes. In particular, bacterial surface proteins present an opportunity for pharmacological intervention and vaccine development. We present Augur, an automatic prediction pipeline that integrates major surface prediction algorithms and enables comparative analysis, classification and visualization for gram-positive bacteria on a genomic scale. AVAILABILITY: http://bioinfo.mikrobio.med.uni-giessen.de/augur

Algorithms↗

In vivo functional proteomics: mammalian genome annotation using CD-tagging.

A self-inactivating CD-tagging retroviral vector was used to introduce epitope and GFP tags into genes and proteins in NIH 3T3 cells. Several hundred cell clones, each expressing GFP fluorescence in a distinctive pattern, were isolated. Molecular analysis showed that a wide variety of genes and proteins, some known and some newly discovered, had been tagged. The analysis also revealed that, in the great majority of instances, the abundance and cellular location of the tagged protein mirrored that of its untagged counterpart. This approach provides a systematic means for the functional annotation of mammalian genomes and proteomes in living cells.

3T3 Cells↗