Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,009 records · Page 56Linked to original sources

A novel sequence similarity searching and visualization method based on overlappingly translated nucleic acids: the blastNP.

Sequence data are stored in nucleic acid and protein databases. Searching the nucleic acid databases is very specific but rather insensitive method. Searching protein databases is sensitive but not very specific procedure. It was expected that the combination of these methods might provide an optimal approach. Therefore an alternative method to TblastX has been developed, known as blastNP. Nucleic acids in database and query sequences were translated into overlapping protein-like sequences (overlappingly translated sequences or OTSs) before searching with blastP. Thus, each nucleic acid sequence is represented by a single "protein like" sequence (instead of three hypothetical proteins in different reading frames). The blastNP method is defined as a blastP that is performed on an overlappingly translated nucleic acid database using a similarly converted nucleic acid query. The specificity and sensitivity of blastNP and TblastX is very similar, however blastNP is more sensitive to detect short sequence similarities (less than 50 residues). BlastNP combines the advantages of nucleotide and protein blasts and bypasses many difficulties: (1). it is more sensitive to weak sequence similarities than blastN, (2). codon redundancy is eliminated, (3). the sensitivity to single nucleotide polymorphism, mutation and sequencing errors are reduced, (4). it is insensitive to frame shifts. This novel method was proved to find significant sequence similarities which remained hidden for other methods and is a promising tool for further understanding (and annotating) the function of many old and new sequences.

Amino Acid Sequence↗

Unraveling the genomic blueprint of the Indian black soldier fly: From genome assembly to evolutionary insights.

The black soldier fly (BSF) (Hermetia illucens) has been renowned for its sustainable bioconversion capabilities, resulting in smart protein production with wide applications in animal feed, bioenergy, and biofertilizer. However, the genetic mechanisms underlying efficient bioconversion and productivity remain poorly understood. To advance strain-specific applications and strengthen genetic resource availability, we present the whole genome sequencing (WGS) data for an Indian isolate of black soldier fly. The assembled genome was 1.46 Gb with a scaffold N50 of 172.7 Mb, and a GC content of 42.6%. Furthermore, 64.17% of genomic sequences were masked as repeated, and 14,317 protein-coding sequences were identified. Variant analysis against the reference genome identified 34.44 million variants (∼33.25 million SNPs and ∼ 1.18 million INDELs), with the majority (99.3%) classified as MODIFIER, 0.54% as LOW impact, 0.14% as MODERATE, and only 0.003% as HIGH impact. Comparative genomic analysis with other related species revealed expansions of gene families in BSF associated with Immune effector (Antimicrobial peptides (AMPs), Lysozymes, and Peptidoglycan Recognition Protein (PGRP) and Detoxification (cytochrome P450 enzymes). Notably, AMPs in the Indian isolate showed enhanced copy number variation in defensin (27) and PGRP (40) compared to reference BSF, suggesting potential regional adaptations to pathogen exposure. Collectively, this genomic data provides an improved resource for evolutionary studies, functional genomics, and targeted genetic improvement of BSF for sustainable bioconversion applications.

Comparative genomics↗

Mining DNA microarray data using a novel approach based on graph theory.

The recent demonstration that biochemical pathways from diverse organisms are arranged in scale-free, rather than random, systems [Jeong et al., Nature 407 (2000) 651-654], emphasizes the importance of developing methods for the identification of biochemical nexuses--the nodes within biochemical pathways that serve as the major input/output hubs, and therefore represent potentially important targets for modulation. Here we describe a bioinformatics approach that identifies candidate nexuses for biochemical pathways without requiring functional gene annotation; we also provide proof-of-principle experiments to support this technique. This approach, called Nexxus, may lead to the identification of new signal transduction pathways and targets for drug design.

Apoptosis↗

Annotating nucleic acid-binding function based on protein structure.

Many of the targets of structural genomics will be proteins with little or no structural similarity to those currently in the database. Therefore, novel function prediction methods that do not rely on sequence or fold similarity to other known proteins are needed. We present an automated approach to predict nucleic-acid-binding (NA-binding) proteins, specifically DNA-binding proteins. The method is based on characterizing the structural and sequence properties of large, positively charged electrostatic patches on DNA-binding protein surfaces, which typically coincide with the DNA-binding-sites. Using an ensemble of features extracted from these electrostatic patches, we predict DNA-binding proteins with high accuracy. We show that our method does not rely on sequence or structure homology and is capable of predicting proteins of novel-binding motifs and protein structures solved in an unbound state. Our method can also distinguish NA-binding proteins from other proteins that have similar, large positive electrostatic patches on their surfaces, but that do not bind nucleic acids.

Amino Acid Motifs↗

Genomic structure and evolutionary context of the human feline leukemia virus subgroup C receptor (hFLVCR) gene: evidence for block duplications and de novo gene formation within duplicons of the hFLVCR locus.

In this paper we sought to analyze the genomic structure and context of human feline leukemia virus subgroup C receptor (hFLVCR), a human glucarate transporter-like gene at chromosome 1q31, and compare it to that of a paralog (FLVCR14q) at chromosome 14q24. Splicing, polyadenylation, and expression patterns, as estimated by in silico analysis, differed between the two FLVCR genes despite their similar genomic structures, suggesting active and independent evolution of transcriptional and messenger RNA processing patterns after gene duplication. Promoter activity was bi-directional for hFLVCR, but not for its 14q paralog. The upstream 1q transcribed sequences were determined to comprise a novel gene of unknown function, LQK1. Annotation of contigs centered at hFLVCR and FLVCRL14q also revealed highly conserved gene clusters on chromosomes 1 and 14, inferred to result from a duplication. The clusters contained members of the FLVCR, Angel (KIAA0759), JDP, p21SNFT, and TGF- families, as well as two uncharacterized families. The genome-wide locations of both previously recognized and four de novo in silico predicted genes belonging to these seven families were determined. Phylogenetic analyses of these families were consistent with the hypothesis that the 1q/14q duplication occurred early within, or immediately prior to the vertebrate divergence, after the protostome-deuterostome divergence but before the amniote-amphibian divergence.

3' Untranslated Regions↗

Fulfilling the promise: drug discovery in the post-genomic era.

The genomic era has brought with it a basic change in experimentation, enabling researchers to look more comprehensively at biological systems. The sequencing of the human genome coupled with advances in automation and parallelization technologies have afforded a fundamental transformation in the drug target discovery paradigm, towards systematic whole genome and proteome analyses. In conjunction with novel proteomic techniques, genome-wide annotation of function in cellular models is possible. Overlaying data derived from whole genome sequence, expression and functional analysis will facilitate the identification of causal genes in disease and significantly streamline the target validation process. Moreover, several parallel technological advances in small molecule screening have resulted in the development of expeditious and powerful platforms for elucidating inhibitors of protein or pathway function. Conversely, high-throughput and automated systems are currently being used to identify targets of orphan small molecules. The consolidation of these emerging functional genomics and drug discovery technologies promises to reap the fruits of the genomic revolution.

Animals↗

Accurate and scalable identification of functional sites by evolutionary tracing.

A common difficulty in post genomics biology is that large-scale techniques of data collection often strip away information on the biological context of these data. The result is a massive number of disconnected observations on sequence, structure, and function from which underlying patterns and biological meaning are obscured. One solution is to build computational filters that pick out sufficiently few facts, relevant to a query, that their relationship is immediately apparent and experimentally testable. Typically, these filters rely on mathematics and statistics, and on first principles from physics and chemistry. We show here that evolution itself can be used to filter sequence and structure data in order to identify evolutionarily important amino acids. A general property of these residues is that they form clusters in native protein structures and point to regions where mutations have the greatest biological impact. The result is an accurate method of functional site annotation that is scalable for structural proteomics.

Amino Acid Sequence↗

Selection system for genes encoding nuclear-targeted proteins.

Nuclear proteins have essential roles in cell proliferation and differentiation. We have developed a yeast selection system-the nuclear transportation trap (NTT)-to identify genes encoding nuclear transport signals. Both unknown and previously identified nuclear localization signals were identified from a human fetal brain cDNA library. The majority (75%) of the unknown proteins examined were exclusively localized to the nucleus in COS-7 cells. We propose that NTT is an efficient method for isolating cDNAs that encode nuclear targeted proteins that can be applied to the retrieval of novel nuclear proteins and to annotate gene function.

Amino Acid Sequence↗

Genome-wide transcription profile of field- and laboratory-selected dichlorodiphenyltrichloroethane (DDT)-resistant Drosophila.

Genome-wide microarray analysis (Affymetrix array) was used (i) to determine whether only one gene, the cytochrome P450 enzyme Cyp6g1, is differentially transcribed in dichlorodiphenyltrichloroethane (DDT)-resistant vs. -susceptible Drosophila; and (ii) to profile common genes differentially transcribed across a DDT-resistant field isolate [Rst(2)DDT(Wisconsin)] and a laboratory DDT-selected population [Rst(2)DDT(91-R)]. Statistical analysis (ANOVA model) identified 158 probe sets that were differentially transcribed among Rst(2)DDT(91-R), Rst(2)DDT(Wisconsin), and the DDT-susceptible genotype Canton-S (P < 0.01). The cytochrome P450 Cyp6a2 and the diazepam-binding inhibitor gene (Dbi) were over transcribed in the two DDT-resistant genotypes when compared to the wild-type Drosophila, and this difference was significant at the most stringent statistical level, a Bonferroni correction. The list of potential candidates differentially transcribed also includes 63 probe sets for which molecular function ontology annotation of the probe sets did not exist. A total of four genes (Cyp6a2, Dbi, Uhg1, and CG11176) were significantly different (P < 5.6 e(-06)) between Rst(2)DDT(91-R) and Canton-S. Additionally, two probe sets encoding Cyp12d1 and Dbi were significantly different between Rst(2)DDT(Wisconsin) and Canton-S after a Bonferroni correction. Fifty-two probe sets, including those associated with pesticide detoxification, ion transport, signal transduction, RNA transcription, and lipid metabolism, were commonly expressed in both resistant lines but were differentially transcribed in Canton-S. Our results suggest that more than Cyp6g1 is overtranscribed in field and laboratory DDT-resistant genotypes, and the number of commonalities suggests that similar resistance mechanisms may exist between laboratory- and field-selected DDT-resistant fly lines.

Animals↗

OpTiles: an R package for adaptive tiling and methylation variability profiling.

SUMMARY: OpTiles is an R package that dynamically defines tiling windows based on the distribution of sequenced CpGs, addressing the limitations of traditional fixed-tiling approaches in targeted methylation datasets. By integrating CpG density with intra-region methylation variability, it provides a reliability metric and extended functionality for annotating, prioritizing, and interpreting complex methylation data. AVAILABILITY AND IMPLEMENTATION: OpTiles is implemented in R and source code is freely available at https://github.com/fhaive/OpTiles. Data are available on Zenodo at https://doi.org/10.5281/zenodo.16961292.

DNA Methylation↗

Essentiality and damage in metabolic networks.

Understanding the architecture of physiological functions from annotated genome sequences is a major task for postgenomic biology. From the annotated genome sequence of the microbe Escherichia coli, we propose a general quantitative definition of enzyme importance in a metabolic network. Using a graph analysis of its metabolism, we relate the extent of the topological damage generated in the metabolic network by the deletion of an enzyme to the experimentally determined viability of the organism in the absence of that enzyme. We show that the network is robust and that the extent of the damage relates to enzyme importance. We predict that a large fraction (91%) of enzymes causes little damage when removed, while a small group (9%) can cause serious damage. Experimental results confirm that this group contains the majority of essential enzymes. The results may reveal a universal property of metabolic networks.

Computer Simulation↗

Create and assess protein networks through molecular characteristics of individual proteins.

MOTIVATION: The study of biological systems, pathways and processes relies increasingly on analyses of networks. Most often, such analyses focus on network topology, thereby treating all proteins or genes as identical, featureless nodes. Integrating molecular data and insights about the qualities of individual proteins into the analysis may enhance our ability to decipher biological pathways and processes. RESULTS: Here, we introduce a novel platform for data integration that generates networks on the macro system-level, analyzes the molecular characteristics of each protein on the micro level, and then combines the two levels by using the molecular characteristics to assess networks. It also annotates the function and subcellular localization of each protein and displays the process on an image of a cell, rendering each protein in its respective cellular compartment. By thus visualizing the network in a cellular context we are able to analyze pathways and processes in a novel way. As an example, we use the system to analyze proteins implicated with Alzheimers disease and show how the integrated view corroborates previous observations and how it helps in the formulation of new hypotheses regarding the molecular underpinnings of the disease. AVAILABILITY: http://www.rostlab.org/services/pinat.

Cell Physiological Phenomena↗

Temporal- and dose-dependent hepatic gene expression changes in immature ovariectomized mice following exposure to ethynyl estradiol.

Temporal- and dose-dependent changes in hepatic gene expression were examined in immature ovariectomized C57BL/6 mice gavaged with ethynyl estradiol (EE), an orally active estrogen. For temporal analysis, mice were gavaged every 24 h for 3 days with 100 microg/kg EE or vehicle and liver samples were collected at 2, 4, 8, 12, 24 and 72 h. Gene expression was monitored using custom cDNA microarrays containing 3067 genes/ESTs of which 393 exhibited a change at one or more time points. Functional gene annotation extracted from public databases associated temporal gene expression changes with growth and proliferation, cytoskeletal and extracellular matrix responses, microtubule-based processes, oxidative metabolism and stress, and lipid metabolism and transport. In the dose-response study, hepatic samples were collected 24 h following treatment with 0, 0.1, 1, 10, 100 or 250 microg/kg EE. Thirty-nine of the 79 genes identified as differentially regulated at 24 h in the time course study exhibited a dose-response relationship with an average ED50 value of 47 +/- 3.5 microg/kg. Comparative analysis indicated that many of the identified temporal and dose-dependent hepatic responses are similar to EE-induced uterine responses reported in the literature and in a companion study using the same animals. Results from these studies confirm that the liver is a highly estrogen responsive tissue that exhibits a number of common responses shared with the uterus as well as distinct estrogen-mediated profiles. These data will further aid in the elucidation of the mechanisms of action of estrogens in the liver as well as in other classical and non-classical estrogen responsive tissues.

Animals↗

Prediction of the coding sequences of unidentified human genes. XXI. The complete sequences of 60 new cDNA clones from brain which code for large proteins.

As an extension of a sequencing project of human cDNA clones which encode large proteins of unidentified genes, we herein present the entire sequences of 60 cDNA clones for the genes named KIAA1879-KIAA1938. The cDNA clones were isolated from size-fractionated cDNA libraries derived from human fetal brain, adult whole brain and amygdala, and their protein-coding sequences were predicted. Thirty-seven cDNA clones entirely sequenced in this study were selected as cDNAs which have coding potentiality by in vitro transcription/translation experiments, and the remaining 23 cDNA clones were chosen by computer-assisted analysis of terminal sequences of cDNAs. The average sizes of the inserts and corresponding open reading frames of cDNA clones analyzed here were 4.5 kb and 2.2 kb (733 amino acid residues), respectively. Sequence analyses against the public databases enabled us to annotate the functions of the predicted products of the 25 genes; 84% of these predicted gene products (21 gene products) were classified into proteins related to cell signaling/communication, nucleic acid management, and cell structure/motility. In addition to the sequence information about these 60 genes, their expression profiles were also studied in some human tissues including brain regions by reverse transcription-coupled polymerase chain reaction, products of which were quantified by enzyme-linked immunosorbent assay.

Adult↗

The SBASE protein domain library, release 7.0: a collection of annotated protein sequence segments.

SBASE 7.0 is the seventh release of the SBASE protein domain library sequences that contains 237 937 annotated structural, functional, ligand-binding and topogenic segments of proteins, cross-referenced to all major sequence databases and sequence pattern collections. The entries are clustered into over 1811 groups and are provided with two WWW-based search facilities for on-line use. SBASE 7.0 is freely available by anonymous 'ftp' file transfer from ftp.icgeb. trieste.it. Automated searching of SBASE with BLAST can be carried out with the WWW servers http://www.icgeb.trieste.it/sbase/and http://sbase.abc.hu/sbase/

Amino Acid Sequence↗

The InterPro database, an integrated documentation resource for protein families, domains and functional sites.

Signature databases are vital tools for identifying distant relationships in novel sequences and hence for inferring protein function. InterPro is an integrated documentation resource for protein families, domains and functional sites, which amalgamates the efforts of the PROSITE, PRINTS, Pfam and ProDom database projects. Each InterPro entry includes a functional description, annotation, literature references and links back to the relevant member database(s). Release 2.0 of InterPro (October 2000) contains over 3000 entries, representing families, domains, repeats and sites of post-translational modification encoded by a total of 6804 different regular expressions, profiles, fingerprints and Hidden Markov Models. Each InterPro entry lists all the matches against SWISS-PROT and TrEMBL (more than 1,000,000 hits from 462,500 proteins in SWISS-PROT and TrEMBL). The database is accessible for text- and sequence-based searches at http://www.ebi.ac.uk/interpro/. Questions can be emailed to interhelp@ebi.ac.uk.

Databases, Factual↗

The SBASE protein domain library, release 8.0: a collection of annotated protein sequence segments.

SBASE 8.0 is the eighth release of the SBASE library of protein domain sequences that contains 294 898 annotated structural, functional, ligand-binding and topogenic segments of proteins, cross-referenced to most major sequence databases and sequence pattern collections. The entries are clustered into over 2005 statistically validated domain groups (SBASE-A) and 595 non-validated groups (SBASE-B), provided with several WWW-based search and browsing facilities for online use. A domain-search facility was developed, based on non-parametric pattern recognition methods, including artificial neural networks. SBASE 8.0 is freely available by anonymous 'ftp' file transfer from ftp.icgeb.trieste.it. Automated searching of SBASE can be carried out with the WWW servers http://www.icgeb.trieste.it/sbase/ and http://sbase.abc. hu/sbase/.

Binding Sites↗

The SBASE protein domain library, release 9.0: an online resource for protein domain identification.

SBASE (http://www.icgeb.trieste.it/sbase) is an online resource of protein domain sequences designed to facilitate detection of domain homologies based on a simple database search. The ninth release of the SBASE library of protein domain sequences contains 320 000 annotated structural, functional, ligand-binding and topogenic segments of proteins clustered into over 3481 domain groups and 483 protein families. Domain identification and functional prediction are based on a comparison of BLAST search outputs with a knowledge base of within-group ('self') and out-of-group ('non-self') similarities of the known domain groups. This is a memory-based approach wherein class-specific similarity functions are automatically learned from the database [Stanfill,C. and Waltz,D. (1986) COMMUN: ACM, 29, 1213-1228].

Animals↗