Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

Gene and protein profiling of the response of MA-10 Leydig tumor cells to human chorionic gonadotropin.

Activation of the steroidogenic machinery by peptide hormones involves a number of steps for transmitting signals from the plasma membrane to mitochondria in a spatially and temporally coordinated manner. Although key proteins mediating the hormonal signal have been identified, recent data suggest that the pathway might involve more complex protein-protein and protein-lipid interactions. Genomic and proteomic methods of analysis, namely the Affymetrix Murine Genome U74A v2 GeneChip and the BD PowerBlot Western Array, were used to identify human chorionic gonadotropin (hCG)-induced changes in mRNA and protein of MA-10 Leydig tumor cells that parallel the increase seen in progesterone synthesis. To analyze the massive amount of data that was generated, a comprehensive protein information matrix summarizing the features of each gene or protein, including its known properties, as well as annotations derived by homology-based functional inference, was developed. Of the genes examined by Affymetrix array, approximately 79 were differentially expressed and of gene products examined by PowerBlot, 9 were differentially expressed (above twofold). Changes in the expression of selected transcripts of interest were confirmed using real-time quantitative polymerase chain reaction and immunoblot analyses. Collectively, these results indicate that hormonal regulation of steroidogenesis is a complex phenomenon, involving proteins that participate in various known and novel pathways, which are implicated in transmitting signals from the plasma membrane to mitochondria and nucleus.

Blotting, Western↗

Persistent biases in the amino acid composition of prokaryotic proteins.

Correspondence analysis of 28 proteomes selected to span the entire realm of prokaryotes revealed universal biases in the proteins' amino acid distribution. Integral Inner Membrane Proteins always form an individual cluster, which can then be used to predict protein localisation in unknown proteomes, independently of the organism's biotope or kingdom. Orphan proteins are consistently rich in aromatic residues. Another bias is also ubiquitous: the amino acid composition is driven by the G + C content of the first codon position. An unexpected bias is driven, in many proteomes, by the AAN box of the genetic code, suggesting some functional biochemical relationship between asparagine and lysine. Less-significant biases are driven by the rare amino acids, cysteine and tryptophan. Some allow identification of species-specific functions or localisation such as surface or exported proteins. Errors in genome annotations are also revealed by correspondence analysis, making it useful for quality control and correction.

Amino Acids↗

Temporal evolution of the Arabidopsis oxidative stress response.

We have carried out a detailed analysis of the changes in gene expression levels in Arabidopsis thaliana ecotype Columbia (Col-0) plants during and for 6 h after exposure to ozone (O3) at 350 parts per billion (ppb) for 6 h. This O3 exposure is sufficient to induce a marked transcriptional response and an oxidative burst, but not to cause substantial tissue damage in Col-0 wild-type plants and is within the range encountered in some major metropolitan areas. We have developed analytical and visualization tools to automate the identification of expression profile groups with common gene ontology (GO) annotations based on the sub-cellular localization and function of the proteins encoded by the genes, as well as to automate promoter analysis for such gene groups. We describe application of these methods to identify stress-induced genes whose transcript abundance is likely to be controlled by common regulatory mechanisms and summarized our findings in a temporal model of the stress response.

Arabidopsis↗

Unravelling the genomic potential of sponge-associated Streptomyces sp. BLC 17-3 from Indonesia for mannooligosaccharide production.

This research aims to show the promising capacity of Streptomyces sp. BLC 17-3 to produce high β-mannanase enzymes and generate mannooligosaccharide (MOS) such as mannobiose, mannotriose, mannotetraose and mannopentaose when exposed to mannan polymers. Streptomyces sp. BLC 17-3 was isolated from the sponge (Rhabdastrella globostellata) Put4 obtained from the marine waters of Putus Island in Bitung, North Sulawesi, Indonesia. The characterization results showed that the peak enzyme activity was achieved at 50 mM sodium acetate, 6.0 pH, and 60 °C temperature on the seventh day of production with a value of 155.77 ± 3.21 U/mL. The SDS-PAGE and zymograms also showed that the size of the enzyme molecule was approximately ±34.8-49.1 kDa. Moreover, whole-genome sequencing was conducted to identify the genetic basis of MOS-synthesizing capabilities in the selected strain, followed by functional annotation of genes encoding mannan degradation and associated functions. The results showed an 8,248,862 Mb complete draft genome of the strain which comprised 111 predicted gene models. Gene annotation also provided important information about the location and function of protein-encoding genes. A total of 6 mannan degradation-related genes encoding mannanase-related metabolism were identified and the three-dimensional structures were predicted using AlphaFold 3. This characterization and modeling further enhanced the bioprospecting and development of this strain which exhibited efficient mannose metabolism. The results showed Streptomyces sp. BLC 17-3 as a promising microorganism for the future bioproduction of MOS which were discovered to have the capability of serving as a potential prebiotic substance to enhance digestion and promote health.

Bioprospecting↗

Chromosome-level genome assembly of Triplophysa scleroptera.

Triplophysa scleroptera is an endemic fish species in Qinghai Lake and the upper reaches of the Yellow River. However, studies on conservation and evolutionary genetics were seriously impeded by the absence of a reference genome. Here, by using PacBio HiFi sequencing and Hi-C assembly technology, we assembled a chromosome-level genome of T. scleroptera, with a total length of 660.22 Mb and 99.82% of the sequence anchored to 25 chromosomes. The contig N50 and scaffold N50 were 9.09 Mb and 24.38 Mb, respectively. The evaluation using BUSCO indicated the genome assembly to be 96.40% complete. About 33.41% of the genome consists of repeat elements. We predicted 26,168 protein-coding genes in the genome, and 99.02% of them were functionally annotated. This high-quality reference genome would serve as a valuable genomic resource for advancing evolutionary conservation genetics studies in this species.

Animals↗

Chromosome-level genome assembly of Ceroplastes pseudoceriferus Green, 1935 (Hemiptera: Coccidae).

Soft scales (Hemiptera: Coccidae) are significant polyphagous pests and majority of which are invasive species. The 364.14 Mb chromosome-level genome of Ceroplastes pseudoceriferus was assembled in this work, with a contig N50 length of 6.16 Mb and scafold N50 length of 21.24 Mb. Approximately 99.89% of assembled sequences were anchored into 18 chromosomes with the assistance of Hi-C reads. Furthermore, approximately 53.98% of the genome was composed of repetitive elements. In total, 10,475 protein-coding genes were predicted, of which 9503 (90.72%) genes were functionally annotated. The BUSCO analysis demonstrated the completeness of the genome annotation is 92.54%. This genome represents first high-quality chromosome level assembly of Coccidae, thereby advancing our knowledge of Coccidae insects and developing effective management strategies that protect crops, forests, and natural ecosystems.

Animals↗

Chromosome-level genome assembly of the longhorn beetle Arhopalus rusticus (Coleoptera: Cerambycidae).

The longhorn beetle Arhopalus rusticus (Coleoptera: Cerambycidae) is a widely distributed wood-boring pest of conifers. Here, we assembled a chromosome-level genome of A. rusticus using Illumina, Oxford Nanopore, and Hi-C sequencing technologies. The assembled genome is 1180.40 Mb, with a scaffold N50 of 125.01 Mb, and BUSCO completeness of 93.6%. All contigs were assembled into ten pseudo-chromosomes. The genome contains 69.87% repeat sequences. We identify 18, 377 protein-coding genes in the genome, of which 11,368 were functionally annotated. This genome provides a valuable resource for understanding the ecology, genetics, and evolution of A. rusticus, as well as for controlling wood-boring pests.

Animals↗

Chromosome-level Genome Assembly of the Halophytic Turfgrass Zoysia macrostachya.

Zoysia macrostachya Franch. & Sav. is a halophytic perennial turfgrass in the Poaceae family, commonly found in the coastal regions of Korea, Japan, and East Asia. Z. macrostachya thrives in high-salinity environments, making it an excellent model for studying abiotic stress resilience. In this study, we present a chromosome-level genome assembly of Z. macrostachya, constructed using Oxford Nanopore long reads, Illumina short reads, and Omni-C sequencing data. The assembly spans 329.78 Mb across 20 chromosomes, with a scaffold N50 of 19.24 Mb, and includes complete telomeric sequences at both ends. The assembly showed 97.8% complete BUSCOs, indicating high genome completeness. Repeat element and gene annotation identified 44.03% of the genome as repetitive elements and 33,474 protein-coding genes. The gene annotation showed 97.1% complete BUSCOs and 86.92% functionally characterized genes. Macrosynteny analysis highlighted highly collinear relationships with related species, providing a foundational understanding of the Z. macrostachya genomic structure. This high-quality genome serves as a valuable resource for advancing salinity tolerance research and improving the genetic diversity of Zoysia species.

Genome, Plant↗

Chromosome-level genome assembly of a cosmopolitan marine harmful algal bloom diatom species Chaetoceros socialis (Chaetocerotaceae).

Chaetoceros socialis is a cosmopolitan diatom species that is crucial for maintaining marine ecosystem structure and driving elemental cycles. C. socialis can form harmful algal blooms (HABs) that may cause a negative impact on the marine ecosystems. Whole-genome information for C. socialis is still unavailable, which may hinder more targeted studies on its ecological adaptive responses and evolutionary drivers. To address this gap, we employed cutting-edge genomic technologies including PacBio single-molecule real-time (SMRT) sequencing and high-throughput chromatin conformation capture (Hi-C) to achieve the first chromosome-level genome assembly of C. socialis. The assembled genome is 60.22 Mb in size with a scaffold N50 of 7.81 Mb and has been anchored to eight pseudochromosomes. A total of 13,378 protein-coding genes were predicted, of which 12,069 (90.22%) were functionally annotated. This high-quality genomic resource provides a fundamental data platform for systematically elucidating the ecological adaptation mechanisms of C. socialis.

Chromosomes↗

The apoptosis database.

The apoptosis database is a public resource for researchers and students interested in the molecular biology of apoptosis. The resource provides functional annotation, literature references, diagrams/images, and alternative nomenclatures on a set of proteins having 'apoptotic domains'. These are the distinctive domains that are often, if not exclusively, found in proteins involved in apoptosis. The initial choice of proteins to be included is defined by apoptosis experts and bioinformatics tools. Users can browse through the web accessible lists of domains, proteins containing these domains and their associated homologs. The database can also be searched by sequence homology using basic local alignment search tool, text word matches of the annotation, and identifiers for specific records. The resource is available at http://www.apoptosis-db.org and is updated on a regular basis.

Animals↗

The LIFEdb database in 2006.

LIFEdb (http://www.LIFEdb.de) integrates data from large-scale functional genomics assays and manual cDNA annotation with bioinformatics gene expression and protein analysis. New features of LIFEdb include (i) an updated user interface with enhanced query capabilities, (ii) a configurable output table and the option to download search results in XML, (iii) the integration of data from cell-based screening assays addressing the influence of protein-overexpression on cell proliferation and (iv) the display of the relative expression ('Electronic Northern') of the genes under investigation using curated gene expression ontology information. LIFEdb enables researchers to systematically select and characterize genes and proteins of interest, and presents data and information via its user-friendly web-based interface.

Cell Proliferation↗

Comparative and functional genomic analysis of prokaryotic nickel and cobalt uptake transporters: evidence for a novel group of ATP-binding cassette transporters.

The transition metals nickel and cobalt, essential components of many enzymes, are taken up by specific transport systems of several different types. We integrated in silico and in vivo methods for the analysis of various protein families containing both nickel and cobalt transport systems in prokaryotes. For functional annotation of genes, we used two comparative genomic approaches: identification of regulatory signals and analysis of the genomic positions of genes encoding candidate nickel/cobalt transporters. The nickel-responsive repressor NikR regulates many nickel uptake systems, though the NikR-binding signal is divergent in various taxonomic groups of bacteria and archaea. B(12) riboswitches regulate most of the candidate cobalt transporters in bacteria. The nickel/cobalt transporter genes are often colocalized with genes for nickel-dependent or coenzyme B(12) biosynthesis enzymes. Nickel/cobalt transporters of different families, including the previously known NiCoT, UreH, and HupE/UreJ families of secondary systems and the NikABCDE ABC-type transporters, showed a mosaic distribution in prokaryotic genomes. In silico analyses identified CbiMNQO and NikMNQO as the most widespread groups of microbial transporters for cobalt and nickel ions. These unusual uptake systems contain an ABC protein (CbiO or NikO) but lack an extracytoplasmic solute-binding protein. Experimental analysis confirmed metal transport activity for three members of this family and demonstrated significant activity for a basic module (CbiMN) of the Salmonella enterica serovar Typhimurium transporter.

ATP-Binding Cassette Transporters↗

Recognition of human genes by stochastic parsing.

A gene finding system, GeneDecoder, based on a parsing technique using a stochastic grammar and dictionary of genetic words is introduced. The structure of human genes are expressed by a stochastic grammar and a dictionary, whose components are the genetic words consisting of genetic phonemes, built as hidden Markov models (HMMs). The HMMs represent the nucleotide acid bases, the codons, and the amino acids. The genetic words in the dictionary are described by the sequence of these HMMs and represent exons, introns, intergenic regions, tRNA regions and signals in DNA sequences. The statistics between these regions are expressed by the grammar, which is a stochastic network of the genetic words. Using the same kind of technique of speech recognition by HMMs with a word dictionary and a grammar, the stochastic network of genetic words enables the motif dictionary to be used during the parsing of the DNA sequences. At the same time, stochastic features of donor/acceptor sites, information of the di-codon statistics, and other important features are integrated into stochastic scores during the parsing. As a result, while the system parses DNA sequences and finds the exon/intron structures, the protein motifs are automatically annotated in the regions. It helps to identify the functions of the genes and reduces the cost of homology search for each hypothetical coding regions. This method is different from simply using the information of homology search. This method uses the information of the motif patterns during the parsing process, but searching the motif patterns after/before finding the coding regions cannot directly affect the parsing process itself. Experimental results have shown that this method reasonably finds and annotates the motifs in the exons in the DNA sequence of human.

Amino Acid Sequence↗

Genome-wide identification of Arabidopsis coiled-coil proteins and establishment of the ARABI-COIL database.

Increasing evidence demonstrates the importance of long coiled-coil proteins for the spatial organization of cellular processes. Although several protein classes with long coiled-coil domains have been studied in animals and yeast, our knowledge about plant long coiled-coil proteins is very limited. The repeat nature of the coiled-coil sequence motif often prevents the simple identification of homologs of animal coiled-coil proteins by generic sequence similarity searches. As a consequence, counterparts of many animal proteins with long coiled-coil domains, like lamins, golgins, or microtubule organization center components, have not been identified yet in plants. Here, all Arabidopsis proteins predicted to contain long stretches of coiled-coil domains were identified by applying the algorithm MultiCoil to a genome-wide screen. A searchable protein database, ARABI-COIL (http://www.coiled-coil.org/arabidopsis), was established that integrates information on number, size, and position of predicted coiled-coil domains with subcellular localization signals, transmembrane domains, and available functional annotations. ARABI-COIL serves as a tool to sort and browse Arabidopsis long coiled-coil proteins to facilitate the identification and selection of candidate proteins of potential interest for specific research areas. Using the database, candidate proteins were identified for Arabidopsis membrane-bound, nuclear, and organellar long coiled-coil proteins.

Algorithms↗

A draft annotation and overview of the human genome.

BACKGROUND: The recent draft assembly of the human genome provides a unified basis for describing genomic structure and function. The draft is sufficiently accurate to provide useful annotation, enabling direct observations of previously inferred biological phenomena. RESULTS: We report here a functionally annotated human gene index placed directly on the genome. The index is based on the integration of public transcript, protein, and mapping information, supplemented with computational prediction. We describe numerous global features of the genome and examine the relationship of various genetic maps with the assembly. In addition, initial sequence analysis reveals highly ordered chromosomal landscapes associated with paralogous gene clusters and distinct functional compartments. Finally, these annotation data were synthesized to produce observations of gene density and number that accord well with historical estimates. Such a global approach had previously been described only for chromosomes 21 and 22, which together account for 2.2% of the genome. CONCLUSIONS: We estimate that the genome contains 65,000-75,000 transcriptional units, with exon sequences comprising 4%. The creation of a comprehensive gene index requires the synthesis of all available computational and experimental evidence.

Chromosome Mapping↗

Predicting functional sites with an automated algorithm suitable for heterogeneous datasets.

BACKGROUND: In a previous report (La et al., Proteins, 2005), we have demonstrated that the identification of phylogenetic motifs, protein sequence fragments conserving the overall familial phylogeny, represent a promising approach for sequence/function annotation. Across a structurally and functionally heterogeneous dataset, phylogenetic motifs have been demonstrated to correspond to a wide variety of functional site archetypes, including those defined by surface loops, active site clefts, and less exposed regions. However, in our original demonstration of the technique, phylogenetic motif identification is dependent upon a manually determined similarity threshold, prohibiting large-scale application of the technique. RESULTS: In this report, we present an algorithmic approach that determines thresholds without human subjectivity. The approach relies on significant raw data preprocessing to improve signal detection. Subsequently, Partition Around Medoids Clustering (PAMC) of the similarity scores assesses sequence fragments where functional annotation remains in question. The accuracy of the approach is confirmed through comparisons to our previous (manual) results and structural analyses. Triosephosphate isomerase and arginyl-tRNA synthetase are discussed as exemplar cases. A quantitative functional site prediction assessment algorithm indicates that the phylogenetic motif predictions, which require sequence information only, are nearly as good as those from evolutionary trace methods that do incorporate structure. CONCLUSION: The automated threshold detection algorithm has been incorporated into MINER, our web-based phylogenetic motif identification server. MINER is freely available on the web at http://www.pmap.csupomona.edu/MINER/. Pre-calculated functional site predictions of the COG database and an implementation of the threshold detection algorithm, in the R statistical language, can also be accessed at the website.

Algorithms↗

Genome-wide screening and functional analysis of protein glycosylation-related genes involved in tomato fruit ripening.

Protein glycosylation, an essential co- and post-translational modification, plays critical roles in plant growth, development, and stress responses. However, its functional role in tomato fruit ripening has not been extensively investigated. Here, key protein glycosylation-related genes involved in tomato fruit ripening were identified by genome-wide screen and subsequently functional characterization. First, a dataset comprising 242 glycosylation-related proteins was established based on Gene Ontology annotations in tomato, combined with sequence homology to protein glycosylation-related proteins from Arabidopsis thaliana and Homo sapiens. Then, Subsequently, 28 genes encoding highly expressed glycosylation-related proteins (RPKM > 30) at the breaker (BR) stage were selected for functional screening, and subsequently 6 genes were identified as regulators of fruit ripening by method of virus-induced gene silencing (VIGS). Among them, Solyc03g098600 (STT3B), Solyc01g109410 (OST48), Solyc04g082670 (RPN1), and Solyc08g076460 (DAD1) functioned as positive regulators of tomato fruit ripening, whereas Solyc04g005340 (UAM2) and Solyc08g075340 (XEG113), acted as negative regulators. The expression of these genes responded dynamically to multiple ripening-related cues, including temperature, light, ethylene, and transcription factors. Furthermore, silencing of these genes individually affected the expression of genes involved in fruit ripening, including ethylene biosynthesis genes (ACS2, ACS4, ACO1, and ACO3), ripening-associated transcription factors (RIN, NOR, NOR-LIKE1, FUL1, and FUL2), and the key gene (PSY1) of lycopene biosynthesis pathway. Collectively, these findings demonstrate that protein glycosylation plays an important role in tomato fruit ripening by modulating ethylene signaling, ripening-associated transcriptional regulation, and lycopene biosynthesis.

Fruit ripening↗

LIFEdb: a database for functional genomics experiments integrating information from external sources, and serving as a sample tracking system.

We have implemented LIFEdb (http://www.dkfz.de/LIFEdb) to link information regarding novel human full-length cDNAs generated and sequenced by the German cDNA Consortium with functional information on the encoded proteins produced in functional genomics and proteomics approaches. The database also serves as a sample-tracking system to manage the process from cDNA to experimental read-out and data interpretation. A web interface enables the scientific community to explore and visualize features of the annotated cDNAs and ORFs combined with experimental results, and thus helps to unravel new features of proteins with as yet unknown functions.

Automation↗