Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

Human specific loss of olfactory receptor genes.

Olfactory receptor (OR) genes constitute the basis for the sense of smell and are encoded by the largest mammalian gene superfamily of >1,000 genes. In humans, >60% of these are pseudogenes. In contrast, the mouse OR repertoire, although of roughly equal size, contains only approximately 20% pseudogenes. We asked whether the high fraction of nonfunctional OR genes is specific to humans or is a common feature of all primates. To this end, we have compared the sequences of 50 human OR coding regions, regardless of their functional annotations, to those of their putative orthologs in chimpanzees, gorillas, orangutans, and rhesus macaques. We found that humans have accumulated mutations that disrupt OR coding regions roughly 4-fold faster than any other species sampled. As a consequence, the fraction of OR pseudogenes in humans is almost twice as high as in the non-human primates, suggesting a human-specific process of OR gene disruption, likely due to a reduced chemosensory dependence relative to apes.

Animals↗

A single lentiviral vector platform for microRNA-based conditional RNA interference and coordinated transgene expression.

RNAi is proving to be a powerful experimental tool for the functional annotation of mammalian genomes. The full potential of this technology will be realized through development of approaches permitting regulated manipulation of endogenous gene expression with coordinated reexpression of exogenous transgenes. We describe the development of a lentiviral vector platform, pSLIK (single lentivector for inducible knockdown), which permits tetracycline-regulated expression of microRNA-like short hairpin RNAs from a single viral infection of any naïve cell system. In mouse embryonic fibroblasts, the pSLIK platform was used to conditionally deplete the expression of the heterotrimeric G proteins Galpha12 and Galpha13 both singly and in combination, demonstrating the Galpha13 dependence of serum response element-mediated transcription. In RAW264.7 macrophages, regulated knockdown of Gbeta2 correlated with a reduced Ca(2+) response to C5a. Insertion of a GFP transgene upstream of the Gbeta2 microRNA-like short hairpin RNA allowed concomitant reexpression of a heterologous mRNA during tetracycline-dependent target gene knockdown, significantly enhancing the experimental applicability of the pSLIK system.

Animals↗

The dual nature of human extracellular superoxide dismutase: one sequence and two structures.

Human extracellular superoxide dismutase (EC-SOD; EC 1.15.1.1) is a scavenger of superoxide anions in the extracellular space. The amino acid sequence is homologous to the intracellular counterpart, Cu/Zn superoxide dismutase (Cu/Zn-SOD), apart from N- and C-terminal extensions. Cu/Zn-SOD is a homodimer containing four cysteine residues within each subunit, and EC-SOD is a tetramer composed of two disulfide-bonded dimers in which each subunit contains six cysteines. The amino acid sequences of all EC-SOD subunits are identical. It is known that Cys-219 is involved in an interchain disulfide. To account for the remaining five cysteine residues we purified human EC-SOD and determined the disulfide bridge pattern. The results show that human EC-SOD exists in two forms, each with a unique disulfide bridge pattern. One form (active EC-SOD) is enzymatically active and contains a disulfide bridge pattern similar to Cu/Zn-SOD. The other form (inactive EC-SOD) has a different disulfide bridge pattern and is enzymatically inactive. The EC-SOD polypeptide chain apparently folds in two different ways, most likely resulting in different three-dimensional structures. Our study shows that one gene may produce proteins with different disulfide bridge arrangements and, thus, by definition, different primary structures. This observation adds another dimension to the functional annotation of the proteome.

Aorta↗

A systematic approach to reconstructing transcription networks in Saccharomycescerevisiae.

Decomposing regulatory networks into functional modules is a first step toward deciphering the logical structure of complex networks. We propose a systematic approach to reconstructing transcription modules (defined by a transcription factor and its target genes) and identifying conditionsperturbations under which a particular transcription module is activateddeactivated. Our approach integrates information from regulatory sequences, genome-wide mRNA expression data, and functional annotation. We systematically analyzed gene expression profiling experiments in which the yeast cell was subjected to various environmental or genetic perturbations. We were able to construct transcription modules with high specificity and sensitivity for many transcription factors, and predict the activation of these modules under anticipated as well as unexpected conditions. These findings generate testable hypotheses when combined with existing knowledge on signaling pathways and protein-protein interactions. Correlating the activation of a module to a specific perturbation predicts links in the cell's regulatory networks, and examining coactivated modules suggests specific instances of crosstalk between regulatory pathways.

Gene Expression Profiling↗

Genome-based identification and analysis of collagen-related structural motifs in bacterial and viral proteins.

Collagens are extended trimeric proteins composed of the repetitive sequence glycine-X-Y. A collagen-related structural motif (CSM) containing glycine-X-Y repeats is also found in numerous proteins often referred to as collagen-like proteins. Little is known about CSMs in bacteria and viruses, but the occurrence of such motifs has recently been demonstrated. Moreover, bacterial CSMs form collagen-like trimers, even though these organisms cannot synthesize hydroxyproline, a critical residue for the stability of the collagen triple helix. Here we present 100 novel proteins of bacteria and viruses (including bacteriophages) containing CSMs identified by in silico analyses of genomic sequences. These CSMs differ significantly from human collagens in amino acid content and distribution; bacterial and viral CSMs have a lower proline content and a preference for proline in the X position of GXY triplets. Moreover, the CSMs identified contained more threonine than collagens, and in 17 of 53 bacterial CSMs threonine was the dominating amino acid in the Y position. Molecular modeling suggests that threonines in the Y position make direct hydrogen bonds to neighboring backbone carbonyls and thus substitute for hydroxyproline in the stabilization of the collagen-like triple-helix of bacterial CSMs. The majority of the remaining CSMs were either rich in proline or rich in charged residues. The bacterial proteins containing a CSM that could be functionally annotated were either surface structures or spore components, whereas the viral proteins generally could be annotated as structural components of the viral particle. The limited occurrence of CSMs in eubacteria and lower eukaryotes and the absence of CSMs in archaebacteria suggests that DNA encoding CSMs has been transferred horizontally, possibly from multicellular organisms to bacteria.

Amino Acid Motifs↗

Proteome analysis of the rice etioplast: metabolic and regulatory networks and novel protein functions.

We report an extensive proteome analysis of rice etioplasts, which were highly purified from dark-grown leaves by a novel protocol using Nycodenz density gradient centrifugation. Comparative protein profiling of different cell compartments from leaf tissue demonstrated the purity of the etioplast preparation by the absence of diagnostic marker proteins of other cell compartments. Systematic analysis of the etioplast proteome identified 240 unique proteins that provide new insights into heterotrophic plant metabolism and control of gene expression. They include several new proteins that were not previously known to localize to plastids. The etioplast proteins were compared with proteomes from Arabidopsis chloroplasts and plastid from tobacco Bright Yellow 2 cells. Together with computational structure analyses of proteins without functional annotations, this comparative proteome analysis revealed novel etioplast-specific proteins. These include components of the plastid gene expression machinery such as two RNA helicases, an RNase II-like hydrolytic exonuclease, and a site 2 protease-like metalloprotease all of which were not known previously to localize to the plastid and are indicative for so far unknown regulatory mechanisms of plastid gene expression. All etioplast protein identifications and related data were integrated into a data base that is freely available upon request.

Amino Acid Sequence↗

The identification of nucleic acid-interacting proteins using a simple proteomics-based approach that directly incorporates the electrophoretic mobility shift assay.

Proteins that interact with nucleic acids are central to numerous cellular processes, and their continuing characterization represents one of the foremost challenges in the postgenomic era. Here we describe a simple proteomics-based approach for the identification by mass spectrometry of proteins in crude extracts that interact with nucleic acids. It incorporates the electrophoretic mobility shift assay and is based on the finding that when a protein forms a complex with nucleic acid its electrophoretic mobility is affected as well as that of the nucleic acid. Our method should greatly reduce and in some cases may even eliminate the need for extensive protein purification and as such should contribute significantly to the functional annotation of the proteome. Furthermore it requires no prior knowledge of the molecular mass, quaternary structure, or pI of the interacting protein. Proof of principle is demonstrated using a recently discovered transcription factor; however, the approach should also have application in the identification of proteins that interact with RNA.

Amino Acid Sequence↗

Discussion on the mechanism of Lingguizhugan Decoction in treating hypertension based on network pharmacology and molecular simulation technology.

To explore the mechanism of Lingguizhugan Decoction in treating hypertension based on network pharmacology and molecular simulation. The active ingredients and potential targets were screened by the Systematic Pharmacological Analysis Platform of Traditional Chinese Medicine (TCMSP). Hypertension-related targets were obtained from OMIM and GeneCards databases. Common targets between drug and hypertension were screened in the Venny platform. A protein-protein interaction (PPI) network was constructed in the STRING database using intersection targets. Key targets in PPI network were analyzed by Cytoscape. R language program was used for Gene Ontology (GO) functional annotation and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analysis. Finally, the binding abilities of the main active ingredients to critical targets were verified by molecular simulation. Naringenin, quercetin, kaempferol, and β-sitosterol in Lingguizhugan Decoction, and potential targets such as STAT3, AKT1, TNF, IL6, JUN, PTGS2, MMP9, CASP3, TP53, and MAPK3, were screened out. KEGG Enrichment analysis revealed that the common targets of Lingguizhugan Decoction and hypertension are mainly involved in the lipid and atherosclerosis signaling pathway, AGE-RAGE signaling pathway in diabetic complications, fluid shear stress and atherosclerosis, and IL17 signaling pathway. The molecular simulation results showed that naringenin-MAPK3, quercetin-MMP9, quercetin-PTGS2, and quercetin-TP53 were the top four in the docking scores. Naringenin-MAPK3 and quercetin-MMP9 were stable, with binding free energies of -27.97 ± 1.41 kcal/mol and -21.15 ± 3.17 kcal/mol, respectively. The possible mechanism of Lingguizhugan Decoction in treating hypertension is characterized of multi-component, multi-target, and multi-pathway.Communicated by Ramaswamy H. Sarma.

Network Pharmacology↗

Exploring biosynthetic potential of the endophytic Penicillium turbatum BLH34 using whole-genome sequence analysis and molecular networking.

An in-depth genomic and metabolomic investigation was conducted on the endophytic fungus Penicillium turbatum BLH34, isolated from Macleaya cordata. Hybrid sequencing (Illumina-Nanopore) generated a high-quality 27.9 Mb genome (GC 48.6%) encoding 9798 proteins, with functional annotation linking 5350 genes to the NCBI non-redundant database and 3404 to KEGG pathways. AntiSMASH analysis uncovered 35 biosynthetic gene clusters (BGCs), 23 of which lacked homology to known pathways, highlighting BLH34's potential for novel metabolite discovery. Molecular networking (GNPS) and LC-MS/MS identified 19 specialised metabolites, including antimicrobial polyketides. Bioassays demonstrated potent inhibition against Staphylococcus aureus (36 mm), Bacillus subtilis (28 mm) and Escherichia coli (24 mm), underscoring its pharmaceutical relevance.

Penicillium↗

Blood-based DNA methylation markers for autism spectrum disorder identification using machine learning.

BACKGROUND: Autism spectrum disorder (ASD) is a complex neurodevelopmental disorder lacking objective biomarkers for early diagnosis. DNA methylation is a promising epigenetic marker, and machine learning offers a data-driven classification approach. However, few studies have examined whole-blood, genome-wide DNA methylation profiles for ASD diagnosis in school-aged children. METHODS: We analyzed genome-wide DNA methylation data from GEO dataset GSE113967, including 52 children with ASD and 48 typically developing (TD) controls. Differentially methylated positions (DMPs) were identified, and feature selection was performed using support vector machine-recursive feature elimination with cross-validation (SVM-RFECV). Classification models were developed using random forest (RF), extreme gradient boosting (XGBoost), and decision tree (DT) classifiers. A nomogram visualized feature contributions. RESULTS: A total of 138 DMPs differentiated ASD from TD children. Eleven CpG sites selected by SVM-RFECV formed the basis for model construction. RF and XGBoost achieved the highest accuracy (75%), with DT reaching 70%. Functional annotation indicated enrichment in cell adhesion and immune-related pathways. CONCLUSIONS: This exploratory study demonstrates the feasibility of integrating peripheral blood DNA methylation data with machine learning to distinguish children with ASD. While limited by sample size and moderate accuracy, this study provides methodological insights into the feasibility of integrating epigenetic and computational approaches for ASD-related biomarker exploration.

Humans↗

Changes of DNA methylation and gene expression profile in placental villi and chorioamniotic membranes under preeclampsia.

BACKGROUND: Preeclampsia (PE) is a serious pregnancy complication with elusive pathogenesis. Although epigenetic dysregulation is implicated, its layer-specific placental roles are poorly defined. This study aimed to identify shared and layer-specific epigenetic alterations in PE by profiling DNA methylation and gene expression in placental villi (PV) and chorioamniotic membranes (CAM). RESEARCH DESIGN AND METHODS: PV and CAM samples were collected from 7 normal and 8 PE pregnancies, and three public DNA methylation datasets (GSE98224, GSE44667, GSE75196) were integrated. Differentially methylated genes (DMGs) and differentially expressed genes (DEGs) were identified based on whole-genome methylation and transcriptome sequencing. Layer-specific and shared gene sets were identified by cross-analysis, with functional annotation using Gene Ontology (GO). RESULTS: EM-seq revealed a hypermethylation-dominant, tissue-specific methylation landscape in PE placentas. Cross-tissue comparison identified shared DMGs between the two layers, including nine key genes consistently altered in public datasets. Integrated analysis in PV further identified 22 co-dysregulated genes, enriched in thermoregulation, maternal-fetal immunity, signal transduction, and cell differentiation. CONCLUSIONS: This study elucidates the shared and layer-specific dysregulation of gene networks at methylomic and transcriptomic levels in PE placenta. Comparing PV and CAM highlights placental epigenetic heterogeneity and dysfunction, offering novel clues for mechanistic research and layer-targeted therapies.

Humans↗

A computational study of Shewanella oneidensis MR-1: structural prediction and functional inference of hypothetical proteins.

The genomes of many organisms have been sequenced in the last 5 years. Typically about 30% of predicted genes from a newly sequenced genome cannot be given functional assignments using sequence comparison methods. In these situations three-dimensional structural predictions combined with a suite of computational tools can suggest possible functions for these hypothetical proteins. Suggesting functions may allow better interpretation of experimental data (e.g., microarray data and mass spectroscopy data) and help experimentalists design new experiments. In this paper, we focus on three hypothetical proteins of Shewanella oneidensis MR-1 that are potentially related to iron transport/metabolism based on microarray experiments. The threading program PROSPECT was used for protein structural predictions and functional annotation, in conjunction with literature search and other computational tools. Computational tools were used to perform transmembrane domain predictions, coiled coil predictions, signal peptide predictions, sub-cellular localization predictions, motif prediction, and operon structure evaluations. Combined computational results from all tools were used to predict roles for the hypothetical proteins. This method, which uses a suite of computational tools that are freely available to academic users, can be used to annotate hypothetical proteins in general.

ATP-Binding Cassette Transporters↗

Biological master games: using biologists' reasoning to guide algorithm development for integrated functional genomics.

We review some powerful new algorithms that build on the intuitive biological interpretation techniques for statistical analysis of functional genomics experiments. Although they were originally designed for transcriptomics, we argue that these algorithms are applicable to any type of -omics study (transcriptomics, proteomics, metabolomics). Rank Products (RP), a strictly non-parametric test statistic to detect differentially regulated elements (genes, proteins, metabolites) in genome-wide screens. RP is particularly powerful for noisy data and low numbers of replicates and makes full use of the availability of a large number of parallel measurements that is typical of modern large-scale experiments. Iterative Group Analysis (iGA), a statistical method that makes the transition from regulated single elements to significant classes of elements, and thus provides an automatic functional annotation of an experiment. Graph-based iGA (GiGA), an extension of iGA that combines experimental data with a broad variety of biological annotations to highlight physiologically relevant regions in a given "evidence graph" (e.g., metabolic networks, signaling pathway diagrams, protein interaction maps). The sequential application of these techniques yields an increasingly abstract interpretation of experimental data that is at the same time quantitative, statistically rigorous, and biologically significant. The results can be used either as helpful tools to guide data visualization and exploration, or as the input for downstream computational applications in a systems biology framework.

Algorithms↗

Analysis and comparison of metabolic pathway databases.

Enormous amounts of data result from genome sequencing projects and new experimental methods. Within this tremendous amount of genomic data 30-40 per cent of the genes being identified in an organism remain unknown in terms of their biological function. As a consequence of this lack of information the overall schema of all the biological functions occurring in a specific organism cannot be properly represented. To understand the functional properties of the genomic data more experimental data must be collected. A pathway database is an effort to handle the current knowledge of biochemical pathways and in addition can be used for interpretation of sequence data. Some of the existing pathway databases can be interpreted as detailed functional annotations of genomes because they are tightly integrated with genomic information. However, experimental data are often lacking in these databases. This paper summarises a list of pathway databases and some of their corresponding biological databases, and also focuses on information about the content and the structure of these databases, the organisation of the data and the reliability of stored information from a biological point of view. Moreover, information about the representation of the pathway data and tools to work with the data are given. Advantages and disadvantages of the analysed databases are pointed out, and an overview to biological scientists on how to use these pathway databases is given.

Animals↗

Delineation of modular proteins: domain boundary prediction from sequence information.

The delineation of domain boundaries of a given sequence in the absence of known 3D structures or detectable sequence homology to known domains benefits many areas in protein science, such as protein engineering, protein 3D structure determination and protein structure prediction. With the exponential growth of newly determined sequences, our ability to predict domain boundaries rapidly and accurately from sequence information alone is both essential and critical from the viewpoint of gene function annotation. Anyone attempting to predict domain boundaries for a single protein sequence is invariably confronted with a plethora of databases that contain boundary information available from the internet and a variety of methods for domain boundary prediction. How are these derived and how well do they work? What definition of 'domain' do they use? We will first clarify the different definitions of protein domains, and then describe the available public databases with domain boundary information. Finally, we will review existing domain boundary prediction methods and discuss their strengths and weaknesses.

Algorithms↗

Unsupervised pattern recognition: an introduction to the whys and wherefores of clustering microarray data.

Clustering has become an integral part of microarray data analysis and interpretation. The algorithmic basis of clustering -- the application of unsupervised machine-learning techniques to identify the patterns inherent in a data set -- is well established. This review discusses the biological motivations for and applications of these techniques to integrating gene expression data with other biological information, such as functional annotation, promoter data and proteomic data.

Algorithms↗

X chromosome-wide association studies for quantitative trait loci based on the mixture of general pedigrees and additional unrelated individuals.

Genome-wide association studies have successfully identified many genetic variants associated with complex traits. However, most existing methods target autosomes rather than X chromosome, and several existing X chromosome-wide association studies (XWAS) at quantitative trait loci (QTL) largely focus on unrelated individuals, with limited attention to general pedigrees or mixture of general pedigrees and additional unrelated individuals (called the mixed data for brevity). In this study, we propose nine novel methods for XWAS at QTL in the mixed data (${\mathrm{MQX}}_{\mathrm{cat}}$, ${\mathrm{MQZ}}_{\mathrm{max}}$, ${\mathrm{MT}}_{\mathrm{plinkw}}$, ${\mathrm{MT}}_{\mathrm{chenw}}$, $\mathrm{MwM}3\mathrm{VNA}$, ${\mathrm{MQMVX}}_{\mathrm{cat}}$, ${\mathrm{MQMVZ}}_{\mathrm{max}}$, $\mathrm{MpMV}$, and $\mathrm{McMV}$), also applicable to general pedigrees alone. The first four methods test for mean differences across genotypes; the latter four test for differences in both means and variances; $\mathrm{MwM}3\mathrm{VNA}$ tests for variance differences only. All mean-based and mean-variance-based methods incorporate X chromosome inactivation information, and all nine methods consider genetic relatedness in pedigrees. Simulation studies confirm well-controlled type I error rates, and inclusion of pedigrees significantly improves statistical power. Note that there has been no study focusing on X chromosome for the mixed data or general pedigrees from UK Biobank database, so we apply our proposed methods to this dataset, which identify five total cholesterol (TC)-associated and 13 low-density lipoprotein cholesterol (LDL-C)-associated single nucleotide polymorphisms (SNPs). Linkage disequilibrium (LD) analysis reveals that these SNPs fall into three distinct LD blocks. Functional annotation and gene ontology enrichment analysis reveal 16 and 28 enriched pathways for TC-associated and LDL-C-associated genes, respectively. These methods provide robust and powerful tools for XWAS at QTL in both mixed data and general pedigrees.

Quantitative Trait Loci↗

Markovian domain fingerprinting: statistical segmentation of protein sequences.

MOTIVATION: Characterization of a protein family by its distinct sequence domains is crucial for functional annotation and correct classification of newly discovered proteins. Conventional Multiple Sequence Alignment (MSA) based methods find difficulties when faced with heterogeneous groups of proteins. However, even many families of proteins that do share a common domain contain instances of several other domains, without any common underlying linear ordering. Ignoring this modularity may lead to poor or even false classification results. An automated method that can analyze a group of proteins into the sequence domains it contains is therefore highly desirable. RESULTS: We apply a novel method to the problem of protein domain detection. The method takes as input an unaligned group of protein sequences. It segments them and clusters the segments into groups sharing the same underlying statistics. A Variable Memory Markov (VMM) model is built using a Prediction Suffix Tree (PST) data structure for each group of segments. Refinement is achieved by letting the PSTs compete over the segments, and a deterministic annealing framework infers the number of underlying PST models while avoiding many inferior solutions. We show that regions of similar statistics correlate well with protein sequence domains, by matching a unique signature to each domain. This is done in a fully automated manner, and does not require or attempt an MSA. Several representative cases are analyzed. We identify a protein fusion event, refine an HMM superfamily classification into the underlying families the HMM cannot separate, and detect all 12 instances of a short domain in a group of 396 sequences. CONTACT: jill@cs.huji.ac.il; tishby@cs.huji.ac.il.

Algorithms↗