Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,405 records · Page 78Linked to original sources

Genome cluster database. A sequence family analysis platform for Arabidopsis and rice.

The genome-wide protein sequences from Arabidopsis (Arabidopsis thaliana) and rice (Oryza sativa) spp. japonica were clustered into families using sequence similarity and domain-based clustering. The two fundamentally different methods resulted in separate cluster sets with complementary properties to compensate the limitations for accurate family analysis. Functional names for the identified families were assigned with an efficient computational approach that uses the description of the most common molecular function gene ontology node within each cluster. Subsequently, multiple alignments and phylogenetic trees were calculated for the assembled families. All clustering results and their underlying sequences were organized in the Web-accessible Genome Cluster Database (http://bioinfo.ucr.edu/projects/GCD) with rich interactive and user-friendly sequence family mining tools to facilitate the analysis of any given family of interest for the plant science community. An automated clustering pipeline ensures current information for future updates in the annotations of the two genomes and clustering improvements. The analysis allowed the first systematic identification of family and singlet proteins present in both organisms as well as those restricted to one of them. In addition, the established Web resources for mining these data provide a road map for future studies of the composition and structure of protein families between the two species.

Algorithms↗

Proteomic analysis of antigens from Leishmania infantum promastigotes.

Leishmaniasis is a zoonotic disease caused by the species of the genus Leishmania, flagellated protozoa that multiply inside mammalian macrophages and are transmitted by the bite of the sandfly. The disease is widespread and due to the lack of fully effective treatment and vaccination the search for new drugs and immune targets is needed. Proteomics seems to be a suitable strategy because the annotated sequenced genome of L. major is available. Here, we present a high-resolution proteome for L. infantum promastigotes comprising of around 700 spots. Western blot with rabbit hyperimmune serum raised against L. infantum promastiogote extracts and further analysis by MALDI-TOF and MALDI-TOF/TOF MS allowed the identification of various relevant functional antigenic proteins. Major antigenic proteins were identified as propionil carboxilasa, ATPase beta subunit, transketolase, proteasome subunit, succinyl-diaminopimelate desuccinylase, a probable tubulin alpha chain, the full-size heat shock protein 70, and several proteins of unknown function. In addition, one enzyme from the ergosterol biosynthesis pathway (adrenodoxin reductase) and the structural paraflagellar rod protein 3 (PAR3) were found among non-antigenic proteins. This study corroborates the usefulness of proteomics in identifying new proteins with crucial biological functions in Leishmania parasites.

Animals↗

MPSS: an integrated database system for surveying a set of proteins.

SUMMARY: We design and implement an integrated database system called 'multi-protein survey system' (MPSS), which provides a platform to retrieve information about many proteins at a time. This system integrates several important and widely used databases including SwissProt, TrEMBL, PDB and InterPro, plus useful references such as GO and KEGG to other databases. Users may submit a group of protein IDs, entry names, SwissProt/TrEMBL accession numbers or GenBank GIs through MPSS' web interface, and obtain protein annotation information from public databases and pre-computed molecular properties speedily. MPSS can also supply comprehensive information about query proteins, including 3D structures, domains, pathway, gene ontology and visual presentation of mapping to the GO tree and KEGG pathway, to provide an up-to-date view of available knowledge with regard to the structures and molecular functions of proteins under study. AVAILABILITY: MPSS is freely accessible at http://www.scbit.org/mpss/

Database Management Systems↗

The transcriptome of adult female Anopheles darlingi salivary glands.

Anopheles (Nyssorhynchus) darlingi is an important malaria vector in South and Central America; however, little is known about molecular aspects of its biology. Genomic and proteomic analyses were performed on the salivary gland products of Anopheles darlingi. A total of 593 randomly selected, salivary gland-derived cDNAs were sequenced and assembled based on their similarities into 288 clusters. The putative translated proteins were classified into three categories: (S) secretory products, (H) housekeeping products and (U) products with unknown cell location and function. Ninety-three clusters encode putative secreted proteins and several of them, such as an anophelin, a thrombin inhibitor, apyrases and several new members of the D7 protein family, were identified as molecules involved in haematophagy. Sugar-feeding related enzymes (alpha-glucosidases and alpha-amylase) also were found among the secreted salivary products. Ninety-nine clusters encode housekeeping proteins associated with energy metabolism, protein synthesis, signal transduction and other cellular functions. Ninety-seven clusters encode proteins with no similarity with known proteins. Comparison of the sequence divergence of the S and H categories of proteins of An. darlingi and An. gambiae revealed that the salivary proteins are less conserved than the housekeeping proteins, and therefore are changing at a faster evolutionary rate. Tabular and supplementary material containing the cDNA sequences and annotations are available at http://www.ncbi.nlm.nih.gov/projects/Mosquito/A_darlingi_sialome/

Animals↗

Proteome analysis of serovars Typhimurium and Pullorum of Salmonella enterica subspecies I.

BACKGROUND: Salmonella enterica subspecies I includes several closely related serovars which differ in host ranges and ability to cause disease. The basis for the diversity in host range and pathogenic potential of the serovars is not well understood, and it is not known how host-restricted variants appeared and what factors were lost or acquired during adaptations to a specific environment. Differences apparent from the genomic data do not necessarily correspond to functional proteins and more importantly differential regulation of otherwise identical gene content may play a role in the diverse phenotypes of the serovars of Salmonella. RESULTS: In this study a comparative analysis of the cytosolic proteins of serovars Typhimurium and Pullorum was performed using two-dimensional gel electrophoresis and the proteins of interest were identified using mass spectrometry. An annotated reference map was created for serovar Typhimurium containing 233 entries, which included many metabolic enzymes, ribosomal proteins, chaperones and many other proteins characteristic for the growing cell. The comparative analysis of the two serovars revealed a high degree of variation amongst isolates obtained from different sources and, in some cases, the variation was greater between isolates of the same serovar than between isolates with different sero-specificity. However, several serovar-specific proteins, including intermediates in sulphate utilisation and cysteine synthesis, were also found despite the fact that the genes encoding those proteins are present in the genomes of both serovars. CONCLUSION: Current microbial proteomics are generally based on the use of a single reference or type strain of a species. This study has shown the importance of incorporating a large number of strains of a species, as the diversity of the proteome in the microbial population appears to be significantly greater than expected. The characterisation of a diverse selection of strains revealed parts of the proteome of S. enterica that alter their expression while others remain stable and allowed for the identification of serovar-specific factors that have so far remained undetected by other methods.

Bacterial Proteins↗

Modeling proteome networks with range-dependent graphs.

In this paper we consider the problem of characterizing and modeling large-scale protein-protein association networks using a class of range-dependent graphs which possess appropriate small world properties. These graphs may be employed in representing given association network using a maximum likelihood approach. This in turn annotates every observed association with its 'range', representing the tendency for such an association to be transitive. The application of a very rapidly developing field of graph theory to the emerging field of proetemics is novel and allows for a many-to-many relationship between individual proteins and groupings of proteins, which in turn may correspond to distinct functional behavior.

Models, Biological↗

Genomic annotation of 15,809 ESTs identified from pooled early gestation human eyes.

To complement cDNA libraries from the human eye at early gestation and to discover candidate genes associated with early ocular development, we used freshly dissected human eyeballs from week 9-14 of gestation to construct the early human fetal eye cDNA library. A total of 15,809 clones were isolated and sequenced from the unamplified and unnormalized library. We screened 11,246 good-quality ESTs, leading to the identification of 5,534 nonredundant clusters. Among them, 4,010 (72%) genes matched in the human protein database (Ensembl). The remaining 28% (1,524) corresponded to potentially novel or previously unidentified ESTs. We used BLASTX to compare our EST data with eight organisms and found common expression of a high portion of genes: Caenorhabditis briggsae (26%), Caenorhabditis elegans (27%), Anopheles gambiae (37%), Drosophila melanogaster (32%), Danio rerio (42%), Fugu rubripes (49%), Rattus norvegicusvalitus (52%), and Mus musculus (59%). Nevertheless, 48% (2,680 of 5,534) of the genes expressed in the early developing eye were not shared with current NEIBank human eye cDNA data. In addition, eight known retinal disease genes existed in our ESTs. Among them, six (COL11A1, BBS5, PDE6B, OAT, VMD2, and PGK1) were conserved among the genomes of other organisms, indicating that our annotated EST set provides not only a valuable resource for gene discovery and functional genomic analysis but also for phylogenetic analysis. Our foremost early gestation human eye cDNA library could provide detailed comparisons across species to identify physiological functions of genes and to elucidate evolutionary mechanisms.

Animals↗

Defining the mammalian CArGome.

Serum response factor (SRF) binds a 1216-fold degenerate cis element known as the CArG box. CArG boxes are found primarily in muscle- and growth-factor-associated genes although the full spectrum of functional CArG elements in the genome (the CArGome) has yet to be defined. Here we describe a genome-wide screen to further define the functional mammalian CArGome. A computational approach involving comparative genomic analyses of human and mouse orthologous genes uncovered >100 hypothetical SRF-dependent genes, including 10 previously identified SRF targets, harboring a conserved CArG element within 4000 bp of the annotated transcription start site (TSS). We PCR-cloned 89 hypothetical SRF targets and subjected each of them to at least two of several validations including luciferase reporter, gel shift, chromatin immunoprecipitation, and mRNA expression following RNAi knockdown of SRF; 60/89 (67%) of the targets were validated. Interestingly, 26 of the validated SRF target genes encode for cytoskeletal/contractile or adhesion proteins. RNAi knockdown of SRF diminishes expression of several SRF-dependent cytoskeletal genes and elicits an attending perturbation in the cytoarchitecture of both human and rodent cells. These data illustrate the power of integrating existing algorithms to interrogate the genome in a relatively unbiased fashion for cis-regulatory element discovery. In this manner, we have further expanded the mammalian CArGome with the discovery of an array of cyto-contractile genes that coordinate normal cytoskeletal homeostasis. We suggest one function of SRF is that of an ancient master regulator of the actin cytoskeleton.

Animals↗

Genetic overlap between depression and C-reactive protein levels: Evidence from a cross-trait analysis.

Inflammation and depression have been consistently associated, with elevated C-reactive protein (CRP) levels observed in a significant subset of affected individuals. However, the genetic mechanisms underlying this association remain poorly understood. We integrated results from large-scale genome-wide association studies (GWAS) of depression and CRP levels in a cross-trait analysis specifically focusing on identifying horizontally pleiotropic loci. Identified variants were stratified as concordant versus discordant based on their direction of effects on the two traits and followed up using functional annotation, gene set enrichment, and colocalization analyses. We also explored causal relationships using Mendelian Randomization (MR) analysis with extensive sensitivity analyses, including adjustment for body mass index (BMI). We identified 9 novel loci. Functional analyses revealed that concordant loci were enriched in genes linked to immune and inflammatory processes, while discordant loci mostly mapped to metabolic pathways, including lipid regulation. MR provided strong evidence for body mass index driving a causal relationship between the genetic liability of depression on CRP levels. Our findings suggest that the association between depression and CRP levels is partly driven by shared genetic influences, pointing to different biological pathways depending on whether genetic effects are concordant or discordant. These results underscore the importance of considering effect direction when assessing the genetic overlap between depression and inflammatory processes. In addition, they highlight BMI as a key factor in the causal relationship between depression and systemic inflammation.

C-Reactive Protein↗

Enhanced statistics for local alignment of multiple alignments improves prediction of protein function and structure.

MOTIVATION: Improved comparisons of multiple sequence alignments (profiles) with other profiles can identify subtle relationships between protein families and motifs significantly beyond the resolution of sequence-based comparisons. RESULTS: The local alignment of multiple alignments (LAMA) method was modified to estimate alignment score significance by applying a new measure based on Fisher's combining method. To verify the new procedure, we used known protein structures, sequence annotations and cyclical relations consistency analysis (CYRCA) sets of consistently aligned blocks. Using the new significance measure improved the sensitivity of LAMA without altering its selectivity. The program performed better than other profile-to-profile methods (COMPASS and Prof_sim) and a sequence-to-profile method (PSI-BLAST). The testing was large scale and used several parameters, including pseudo-counts profile calculations and local ungapped blocks or more extended gapped profiles. This comparison provides guidelines to the relative advantages of each method for different cases. We demonstrate and discuss the unique advantages of using block multiple alignments of protein motifs.

Algorithms↗

CyanoBase, the genome database for Synechocystis sp. strain PCC6803: status for the year 2000.

CyanoBase provides an online resource for access to data on genomic information about the cyanobacterium Synechocystis sp. strain PCC6803. The database contains annotations for each protein-coding gene deduced from the entire nucleotide sequence of the genome, gene classification lists, and keyword and similarity search engines. Core portions of CyanoBase consist of annotations for each of the 3168 protein genes deduced from the entire nucleotide sequence of this genome. The contents of each gene were improved by updating with the results of similarity searches and by introducing references for analysis in bioinformatics. The database now contains repository facilities that store and provide experimental information, in addition to providing proposals for the function of each gene. This information should help to avoid unnecessary, overlapping experiments and should assist communication between scientists who wish to elucidate the function of putative genes on the cyanobacteria genome. The current URL of CyanoBase is http://www.kazusa.or.jp:8080/cyano/

Cyanobacteria↗

POPSCOMP: an automated interaction analysis of biomolecular complexes.

Large-scale analysis of biomolecular complexes reveals the functional network within the cell. Computational methods are required to extract the essential information from the available data. The POPSCOMP server is designed to calculate the interaction surface between all components of a given complex structure consisting of proteins, DNA or RNA molecules. The server returns matrices and graphs of surface area burial that can be used to automatically annotate components and residues that are involved in complex formation, to pinpoint conformational changes and to estimate molecular interaction energies. The analysis can be performed on a per-atom level or alternatively on a per-residue level for low-resolution structures. Here, we present an analysis of ribosomal structures in complex with various antibiotics to exemplify the potential and limitations of automated complex analysis. The POPSCOMP server is accessible at http://ibivu.cs.vu.nl/programs/popscompwww/.

Anti-Bacterial Agents↗

Identification of putative noncoding RNAs among the RIKEN mouse full-length cDNA collection.

With the sequencing and annotation of genomes and transcriptomes of several eukaryotes, the importance of noncoding RNA (ncRNA)-RNA molecules that are not translated to protein products-has become more evident. A subclass of ncRNA transcripts are encoded by highly regulated, multi-exon, transcriptional units, are processed like typical protein-coding mRNAs and are increasingly implicated in regulation of many cellular functions in eukaryotes. This study describes the identification of candidate functional ncRNAs from among the RIKEN mouse full-length cDNA collection, which contains 60,770 sequences, by using a systematic computational filtering approach. We initially searched for previously reported ncRNAs and found nine murine ncRNAs and homologs of several previously described nonmouse ncRNAs. Through our computational approach to filter artifact-free clones that lack protein coding potential, we extracted 4280 transcripts as the largest-candidate set. Many clones in the set had EST hits, potential CpG islands surrounding the transcription start sites, and homologies with the human genome. This implies that many candidates are indeed transcribed in a regulated manner. Our results demonstrate that ncRNAs are a major functional subclass of processed transcripts in mammals.

Animals↗

Glycosylphosphatidylinositol lipid anchoring of plant proteins. Sensitive prediction from sequence- and genome-wide studies for Arabidopsis and rice.

Posttranslational glycosylphosphatidylinositol (GPI) lipid anchoring is common not only for animal and fungal but also for plant proteins. The attachment of the GPI moiety to the carboxyl-terminus after proteolytic cleavage of a C-terminal propeptide is performed by the transamidase complex. Its four known subunits also have obvious full-length orthologs in the Arabidopsis and rice (Oryza sativa) genomes; thus, the mechanism of substrate protein processing appears similar for all eukaryotes. A learning set of plant proteins (substrates for the transamidase complex) has been collected both from the literature and plant sequence databases. We find that the plant GPI lipid anchor motif differs in minor aspects from the animal signal (e.g. the plant hydrophobic tail region can contain a higher fraction of aromatic residues). We have developed the "big-Pi plant" program for prediction of compatibility of query protein C-termini with the plant GPI lipid anchor motif requirements. Validation tests show that the sensitivity for transamidase targets is approximately 94%, and the rate of false positive prediction is about 0.1%. Thus, the big-Pi predictor can be applied as unsupervised genome annotation and target selection tool. The program is also suited for the design of modified protein constructs to test their GPI lipid anchoring capacity. The big-Pi plant predictor Web server and lists of potential plant precursor proteins in Swiss-Prot, SPTrEMBL, Arabidopsis, and rice proteomes are available at http://mendel.imp.univie.ac.at/gpi/plants/gpi_plants.html. Arabidopsis and rice protein hits have been functionally classified. Several GPI lipid-anchored arabinogalactan-related proteins have been identified in rice.

Arabidopsis↗

[A evolutionary approach to identification of orthologous relationship across proteomes].

How to identify the true orthologous and paralogous relationships among protein families is still a key problem in genome annotation and comparative protemics. Here, a evolutionary approach to ascertainment of the orthologous relationships across the genomes is developed. Forty-four cases of protein families are used in the test for the evolutionary approach. Compared with the method of COG (cluster of orthologous groups of proteins), this approach can generally identify the orthologous relationships and accurately predict the genome function.

Animals↗

CyanoBase, a www database containing the complete nucleotide sequence of the genome of Synechocystis sp. strain PCC6803.

CyanoBase (http://www.kazusa.or.jp/cyano/) is a database containing genomic information on the cyanobacterium Synechocystis sp. strain PCC6803. It furnishes an annotation to each of the 3168 protein genes deduced from the entire nucleotide sequence of this genome. Information on the genome can be directly accessed through three different menus: a clickable physical map of the genome, a gene classification list, and a keyword search menu, all of which are accessible from the main page of the database. The entry page for a gene annotation contains the following information: the location of the gene on the genome, the nucleotide and deduced amino acid sequence of the gene, the result of a similarity search, and the classification of the deduced gene product according to its function. This page has reverse-links to the local physical map and gene classification list so that relevant genes can be searched in terms of their location on the genome and their function. In addition, the main page of CyanoBase provides engines for similarity searches between a query sequence and the entire genome sequence and for keyword searches, in addition to numerous links to pages containing related information.

Bacterial Proteins↗

Assembly and Characterization of the First Complete Mitochondrial Genome of Tussilago farfara L.: Insights into Biological Functions and Phylogenetic Relationships within the Asteraceae Family.

Tussilago farfara L., a member of the Asteraceae family, is an economically valuable species due to its edible and medicinal properties. To elucidate the structural characteristics, genetic mechanisms, and evolutionary pathways of the organelle genomes of T. farfara, we sequenced, assembled, and annotated its mitochondrial genome for the first time. The complete mitochondrial genome of T. farfara spans 306,024 bp and contains 33 mitochondrial protein-coding genes (PCGs), 3 rRNAs, and 22 tRNAs. Analysis of the nucleotide substitution rate and genetic diversity revealed that most mitochondrial genome genes may have undergone purifying selection, indicating a slow evolutionary rate and a relatively conserved genomic structure. We further identified 13 fragments of chloroplast-derived DNA integrated into the mitochondrial genome, evidencing intracellular gene transfer. Collinearity analysis showed that Arctium lappa shares the most extensive mitochondrial homologous sequences and the highest sequence similarity with T. farfara. Phylogenetic analysis based on the mitochondrial genome helped to clarify the evolutionary and taxonomic position of T. farfara within the Asteraceae family. The mitochondrial genome sequence of T. farfara provides a valuable genomic resource for species identification and for evolutionary studies within the Asteraceae family.

Genome, Mitochondrial↗

Differential gene expression analysis during porcine hepatocyte spheroid formation.

Primary porcine hepatocytes cultured in suspension self-assemble into multicellular aggregates or spheroids that display enhanced liver-specific functional capability and remain viable for an extended period of time in vitro. The molecular events underlying the process of spheroid formation were explored by differential gene expression analysis. Critical time points in spheroid formation were first identified by reverse transcriptase-polymerase chain reaction (RT-PCR) analysis of stress-related gene expression levels at different stages of spheroid formation. Suppression subtractive hybridization was used to identify transcripts up- or down-regulated at different stages of spheroid formation. Subsequently, three sets of reciprocal subtractions, comparing freshly isolated hepatocytes, spheroid-forming hepatocytes, and mature spheroids were carried out, and differentially expressed transcripts were isolated, cloned, sequenced, and annotated. A total of 65 genes and 14 novel transcripts were identified as differentially expressed, and very high sequence conservation between pig and human transcripts was observed. The resultant expressed sequence tags (ESTs) revealed a rapid decrease in the transcript levels of a subset of liver-specific genes, cytochrome P450s, and enzymes involved in heme biosynthesis, as well as up-regulation of genes involved in calcium-dependent vesicle trafficking and a number of acute-phase proteins in mature spheroids. Previous morphological and functional data on hepatocyte spheroid formation support cellular polarization of the hepatocyte into apical and basolateral domains in spheroids. This is important for the re-emergence of differentiated functions in vitro and is reflected by differences in gene expression patterns.

Animals↗