Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,405 records · Page 78Linked to original sources

Proteomic analysis of low-abundant integral plasma membrane proteins based on gels.

To characterize low-copy integral membrane proteins and offer some methods for human liver proteome projects, we fractionated highly purified rat liver plasma membrane (PM). PM was purified through two sucrose density gradient centrifugations, and treated with 0.1 M Na(2)CO(3), chloroform/methanol and Triton X-100. Proteins were separated by electrophoresis and submitted to mass spectrometry analysis. Four hundred and fifty-seven non-redundant membrane proteins were identified, of which 23% (105) were integral membrane proteins with one or more transmembrane domains. One hundred and fifty-three (33.5%) had no location annotation and 68 were unknown-function proteins. The proteins from different fractions were complementory. A database search for all identified proteins revealed that 53 proteins were involved in the cell communication pathway. More interestingly, more than 50% of the proteins had a protein abundance index concentration of less than 0.1 mol/l, and 12% proteins a concentration 100 times less than that of arginase 1 and actin.

Animals↗

Assessment of a systematic expression profiling approach in ENU-induced mouse mutant lines.

Comparative genomewide expression profiling is a powerful tool in the effort to annotate the mouse genome with biological function. The systematic analysis of RNA expression data of mouse lines from the Munich ENU mutagenesis screen might support the understanding of the molecular biology of such mutants and provide new insights into mammalian gene function. In a direct comparison of DNA microarray experiments of individual versus pooled RNA samples of organs from ENU-induced mouse mutants, we provide evidence that individual RNA samples may outperform pools in some aspects. Genes with high biological variability in their expression levels (noisy genes) are identified as false positives in pooled samples. Evidence suggests that highly stringent housing conditions and standardized procedures for the isolation of organs significantly reduce biological variability in gene expression profiling experiments. Data on wild-type individuals demonstrate the positive effect of controlling variables such as social status, food intake before organ sampling, and stress with regard to reproducibility of gene expression patterns. Analyses of several organs from various ENU-induced mutant lines in general show low numbers of differentially expressed genes. We demonstrate the feasibility to detect transcriptionally affected organs employing RNA expression profiling as a tool for molecular phenotyping.

Animals↗

Unlocking the molecular engineering of Geobacillus glycoside hydrolases as a source of industrial biocatalysts.

This review examines Geobacillus sensu stricto as a source of thermostable glycoside hydrolases (GH) for biomass conversion, food processing, and enzyme engineering. Recent peer-reviewed literature was assessed with emphasis on taxonomy, genome-based Carbohydrate-Active Enzymes (CAZyme) prediction, biochemical validation, structural data, and engineering case studies. Taxonomic boundaries were interpreted using current Anoxybacillaceae frameworks, with Parageobacillus treated as a related comparator rather than as Geobacillus. The strongest evidence supports GH13 alpha-amylases, xylan-active systems, beta-xylosidases, and selected accessory enzymes. Recent studies also show that genome mining must be coupled with enzymatic assays and product profiling because CAZyme annotation alone does not prove industrial function. Molecular engineering has improved relevant traits, including the longer thermal half-life of engineered G. stearothermophilus alpha-amylase variants, the increased catalytic efficiency of oligo-alpha-1,6-glucosidase variants, and improved AmyS expression in Bacillus subtilis. Geobacillus glycoside hydrolases are best interpreted as process-specific, engineerable biocatalytic templates. Their translation requires reliable taxonomy, functional validation, structural interpretation, scalable expression and testing on realistic substrates. This synthesis also recognises current limitations: many predicted CAZymes still lack biochemical validation, complete cellulolytic systems remain less mature than xylan- and starch-active systems, and scale-up data remain scarce.

Geobacillus↗

Identification of functional modules in a PPI network by clique percolation clustering.

Large-scale experiments and data integration have provided the opportunity to systematically analyze and comprehensively understand the topology of biological networks and biochemical processes in cells. Modular architecture which encompasses groups of genes/proteins involved in elementary biological functional units is a basic form of the organization of interacting proteins. Here we apply a graph clustering algorithm based on clique percolation clustering to detect overlapping network modules of a protein-protein interaction (PPI) network. Our analysis of the yeast Sacchromyces cerevisiae suggests that most of the detected modules correspond to one or more experimentally functional modules and half of these annotated modules match well with experimentally determined protein complexes. Our method of analysis can of course be applied to protein-protein interaction data for any species and even other biological networks.

Algorithms↗

The sea urchin kinome: a first look.

This paper reports a preliminary in silico analysis of the sea urchin kinome. The predicted protein kinases in the sea urchin genome were identified, annotated and classified, according to both function and kinase domain taxonomy. The results show that the sea urchin kinome, consisting of 353 protein kinases, is closer to the Drosophila kinome (239) than the human kinome (518) with respect to total kinase number. However, the diversity of sea urchin kinases is surprisingly similar to humans, since the urchin kinome is missing only 4 of 186 human subfamilies, while Drosophila lacks 24. Thus, the sea urchin kinome combines the simplicity of a non-duplicated genome with the diversity of function and signaling previously considered to be vertebrate-specific. More than half of the sea urchin kinases are involved with signal transduction, and approximately 88% of the signaling kinases are expressed in the developing embryo. These results support the strength of this nonchordate deuterostome as a pivotal developmental and evolutionary model organism.

Animals↗

Conservation patterns in different functional sequence categories of divergent Drosophila species.

We have explored the distributions of fully conserved ungapped blocks in genome-wide pair-wise alignments of recently completed species of Drosophila: D. melanogaster, D. yakuba, D. ananassae, D. pseudoobscura, D. virilis, and D. mojavensis. Based on these distributions we have found that nearly every functional sequence category possesses its own distinctive conservation pattern, sometimes independent of the overall sequence conservation level. In the coding and regulatory regions, the ungapped blocks were longer than in introns, UTRs, and nonfunctional sequences. At the same time, the blocks in the coding regions carried a 3N + 2 signature characteristic of synonymous substitutions in the third-codon position. Larger block sizes in transcription regulatory regions can be explained by the presence of conserved arrays of binding sites for transcription factors. We also have shown that the longest ungapped blocks, or "ultraconserved" sequences, are associated with specific gene groups, including those encoding ion channels and components of the cytoskeleton. We discuss how restraining conservation patterns may help in mapping functional sequence categories and improve genome annotation.

Animals↗

The gene identification problem: an overview for developers.

The gene identification problem is the problem of interpreting nucleotide sequences by computer, in order to provide tentative annotation on the location, structure, and functional class of protein-coding genes. This problem is of self-evident importance, and is far from being fully solved, particularly for higher eukaryotes. Thus it is not surprising that the number of algorithm and software developers working in the area is rapidly increasing. The present paper is an overview of the field, with an emphasis on eukaryotes, for such developers.

Base Sequence↗

mRNA 5' region sequence incompleteness: a potential source of systematic errors in translation initiation codon assignment in human mRNAs.

The amino acid sequence of gene products is routinely deduced from the nucleotide sequence of the relative cloned cDNA, according to the rules for recognition of start codon (first-AUG rule, optimal sequence context) and the genetic code. From this prediction stem most subsequent types of product analysis, although all standard methods for cDNA cloning are affected by a potential inability to effectively clone the 5' region of mRNA. Revision by bioinformatics and cloning methods of 109 known genes located on human chromosome 21 (HC 21) shows that 60 mRNAs lack any in-frame stop upstream of the first-AUG, and that in five cases (DSCR1, KIAA0184, KIAA0539, SON, and TFF3) the coding region at the 5' end was incompletely characterized in the original descriptions. We describe the respective consequences for genomic annotation, domain and ortholog identification, and functional experiments design. We have also analyzed the sequences of 13,124 human mRNAs (RefSeq databank), discovering that in 6448 cases (49%), an in-frame stop codon is present upstream of the initiation codon, while in the other 6676 mRNAs (51%), identification of additional bases at the mRNA 5' region could well reveal some new upstream in-frame AUG codons in the optimal context. Proportionally to the HC 21 data, about 550 known human genes might thus be affected by this 5' end mRNA artifact.

5' Untranslated Regions↗

HMM-based databases in InterPro.

Protein family databases are an important resource for protein annotation and understanding protein evolution and function. In recent years hidden Markov models (HMMs) have become one of the key technologies used for detection of members of these families. This paper reviews the Pfam, TIGRFAMs and SMART databases that use the profile-HMMs provided by the HMMER package.

Computational Biology↗

Java-based application framework for visualization of gene regulatory region annotations.

MOTIVATION: The genome sequences of several organisms are either complete, or being sequenced. Each genome needs to be integrated with various types of annotations, e.g. locations of genes, promoters and other functional elements such as transcriptional regulatory elements. A robust application framework will be useful for developing web-based applications to visualize various genome annotations. RESULTS: We developed genome data visualization toolkit (GDVTK) as an application framework that consists of a set of data structures and core classes, using Java technology. GDVTK is a sound framework for developing web-based applications to present the gene regulatory region annotations in visual form. The current version of GDVTK consists of eight packages and 38 Java classes that are portable, reusable and extensible for plugging in new data sources and models. We implemented GDVTK for visualization of promoter annotations in Mammalian Promoter Database (MPromDb), a web-based gene-regulatory information server. AVAILABILITY: GDVTK is available under GNU general public license. Source code and software documentation can be found at the URL http://bioinformatics.med.ohio-state.edu/GDVTK.

Computer Graphics↗

Aphid biology: expressed genes from alate Toxoptera citricida, the brown citrus aphid.

The brown citrus aphid, Toxoptera citricida (Kirkaldy), is considered the primary vector of citrus tristeza virus, a severe pathogen which causes losses to citrus industries worldwide. The alate (winged) form of this aphid can readily fly long distances with the wind, thus spreading citrus tristeza virus in citrus growing regions. To better understand the biology of the brown citrus aphid and the emergence of genes expressed during wing development, we undertook a large-scale 5' end sequencing project of cDNA clones from alate aphids. Similar large-scale expressed sequence tag (EST) sequencing projects from other insects have provided a vehicle for answering biological questions relating to development and physiology. Although there is a growing database in GenBank of ESTs from insects, most are from Drosophila melanogaster and Anopheles gambiae, with relatively few specifically derived from aphids. However, important morphogenetic processes are exclusively associated with piercing-sucking insect development and sap feeding insect metabolism. In this paper, we describe the first public data set of ESTs from the brown citrus aphid, T. citricida. The cDNA library was derived from alate adults due to their significance in spreading viruses (e.g., citrus tristeza virus). Over 5180 cDNA clones were sequenced, resulting in 4263 high-quality ESTs. Contig alignment of these ESTs resulted in 2124 total assembled sequences, including both contiguous sequences and singlets. Approximately 33% of the ESTs currently have no significant match in either the non-redundant protein or nucleic acid databases. Sequences returning matches with an E-value of < or = -10 using BLASTX, BLASTN, or TBLASTX were annotated based on their putative molecular function and biological process using the Gene Ontology classification system. These data will aid research efforts in the identification of important genes within insects, specifically aphids and other sap feeding insects within the Order Hemiptera.

Animals↗

SMART: a web-based tool for the study of genetically mobile domains.

SMART (a Simple Modular Architecture Research Tool) allows the identification and annotation of genetically mobile domains and the analysis of domain architectures (http://SMART.embl-heidelberg.de ). More than 400 domain families found in signalling, extra-cellular and chromatin-associated proteins are detectable. These domains are extensively annotated with respect to phyletic distributions, functional class, tertiary structures and functionally important residues. Each domain found in a non-redundant protein database as well as search parameters and taxonomic information are stored in a relational database system. User interfaces to this database allow searches for proteins containing specific combinations of domains in defined taxa.

Database Management Systems↗

MIPS: a database for genomes and protein sequences.

The Munich Information Center for Protein Sequences (MIPS-GSF), Martinsried, near Munich, Germany, continues its longstanding tradition to develop and maintain high quality curated genome databases. In addition, efforts have been intensified to cover the wealth of complete genome sequences in a systematic, comprehensive form. Bioinformatics, supporting national as well as European sequencing and functional analysis projects, has resulted in several up-to-date genome-oriented databases. This report describes growing databases reflecting the progress of sequencing the Arabidopsis thaliana (MATDB) and Neurospora crassa genomes (MNCDB), the yeast genome database (MYGD) extended by functional analysis data, the database of annotated human EST-clusters (HIB) and the database of the complete cDNA sequences from the DHGP (German Human Genome Project). It also contains information on the up-to-date database of complete genomes (PEDANT), the classification of protein sequences (ProtFam) and the collection of protein sequence data within the framework of the PIR-International Protein Sequence Database. These databases can be accessed through the MIPS WWW server (http://www. mips.biochem.mpg.de).

Arabidopsis↗

Protein Information Resource: a community resource for expert annotation of protein data.

The Protein Information Resource, in collaboration with the Munich Information Center for Protein Sequences (MIPS) and the Japan International Protein Information Database (JIPID), produces the most comprehensive and expertly annotated protein sequence database in the public domain, the PIR-International Protein Sequence Database. To provide timely and high quality annotation and promote database interoperability, the PIR-International employs rule-based and classification-driven procedures based on controlled vocabulary and standard nomenclature and includes status tags to distinguish experimentally determined from predicted protein features. The database contains about 200,000 non-redundant protein sequences, which are classified into families and superfamilies and their domains and motifs identified. Entries are extensively cross-referenced to other sequence, classification, genome, structure and activity databases. The PIR web site features search engines that use sequence similarity and database annotation to facilitate the analysis and functional identification of proteins. The PIR-Inter-national databases and search tools are accessible on the PIR web site at http://pir.georgetown.edu/ and at the MIPS web site at http://www.mips.biochem.mpg.de. The PIR-International Protein Sequence Database and other files are also available by FTP.

Computational Biology↗

Identification of stop codon readthrough genes in Saccharomyces cerevisiae.

We specifically sought genes within the yeast genome controlled by a non-conventional translation mechanism involving the stop codon. For this reason, we designed a computer program using the yeast database genomic regions, and seeking two adjacent open reading frames separated only by a unique stop codon (called SORFs). Among the 58 SORFs identified, eight displayed a stop codon bypass level ranging from 3 to 25%. For each of the eight sequences, we demonstrated the presence of a poly(A) mRNA. Using isogenic [PSI(+)] and [psi(-)] yeast strains, we showed that for two of the sequences the mechanism used is a bona fide readthrough. However, the six remaining sequences were not sensitive to the PSI state, indicating either a translation termination process independent of eRF3 or a new stop codon bypass mechanism. Our results demonstrate that the presence of a stop codon in a large ORF may not always correspond to a sequencing error, or a pseudogene, but can be a recoding signal in a functional gene. This emphasizes that genome annotation should take into account the fact that recoding signals could be more frequently used than previously expected.

Base Sequence↗

Recent additions and improvements to the Onto-Tools.

The Onto-Tools suite is composed of an annotation database and six seamlessly integrated, web-accessible data mining tools: Onto-Express, Onto-Compare, Onto-Design, Onto-Translate, Onto-Miner and Pathway-Express. The Onto-Tools database has been expanded to include various types of data from 12 new databases. Our database now integrates different types of genomic data from 19 sequence, gene, protein and annotation databases. Additionally, our database is also expanded to include complete Gene Ontology (GO) annotations. Using the enhanced database and GO annotations, Onto-Express now allows functional profiling for 24 organisms and supports 17 different types of input IDs. Onto-Translate is also enhanced to fully utilize the capabilities of the new Onto-Tools database with an ultimate goal of providing the users with a non-redundant and complete mapping from any type of identification system to any other type. Currently, Onto-Translate allows arbitrary mappings between 29 types of IDs. Pathway-Express is a new tool that helps the users find the most interesting pathways for their input list of genes. Onto-Tools are freely available at http://vortex.cs.wayne.edu/Projects.html.

Animals↗

The LIFEdb database in 2006.

LIFEdb (http://www.LIFEdb.de) integrates data from large-scale functional genomics assays and manual cDNA annotation with bioinformatics gene expression and protein analysis. New features of LIFEdb include (i) an updated user interface with enhanced query capabilities, (ii) a configurable output table and the option to download search results in XML, (iii) the integration of data from cell-based screening assays addressing the influence of protein-overexpression on cell proliferation and (iv) the display of the relative expression ('Electronic Northern') of the genes under investigation using curated gene expression ontology information. LIFEdb enables researchers to systematically select and characterize genes and proteins of interest, and presents data and information via its user-friendly web-based interface.

Cell Proliferation↗

Genomic Language Model for Predicting Enhancers and Their Allele-Specific Activity in the Human Genome.

Predicting and deciphering the regulatory logic of enhancers is a challenging problem, due to the intricate sequence features and lack of consistent genetic or epigenetic signatures that can accurately discriminate enhancers from other genomic regions. Recent machine-learning based methods have spotlighted the importance of extracting nucleotide composition of enhancers but failed to learn the sequence context and perform suboptimally. Motivated by advances in genomic language models, we developed DNABERT-Enhancer, a novel enhancer prediction method, by applying DNABERT pre-trained language model on the human genome. We trained two different models, using large collection of enhancers curated from the ENCODE registry of candidate cis-Regulatory Elements. The best fine-tuned model achieved 88.05% accuracy with Matthews correlation coefficient of 76% on independent set aside data. Further, we present the analysis of the predicted enhancers for all chromosomes of the human genome by comparing with the enhancer regions reported in publicly available databases. Finally, we applied DNABERT-Enhancer along with other DNABERT based regulatory genomic region prediction models to predict candidate SNPs with allele-specific enhancer and transcription factor binding activity. The genome-wide enhancer annotations and candidate loss-of-function genetic variants predicted by DNABERT-Enhancer provide valuable resources for genome interpretation in functional and clinical genomics studies.

Journal Article↗