Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,135 records · Page 63Linked to original sources

Identifying the 3'-terminal exon in human DNA.

MOTIVATION: We present JTEF, a new program for finding 3' terminal exons in human DNA sequences. This program is based on quadratic discriminant analysis, a standard non-linear statistical pattern recognition method. The quadratic discriminant functions used for building the algorithm were trained on a set of 3' terminal exons of type 3tuexon (those containing the true STOP codon). RESULTS: We showed that the average predictive accuracy of JTEF is higher than the presently available best programs (GenScan and Genemark.hmm) based on a test set of 65 human DNA sequences with 121 genes. In particular JTEF performs well on larger genomic contigs containing multiple genes and significant amounts of intergenic DNA. It will become a valuable tool for genome annotation and gene functional studies. AVAILABILITY: JTEF is available free for academic users on request from ftp://cshl.org/pub/science/mzhanglab/JTEF and will be made available through the World Wide Web (http://argon.cshl.org/).

Algorithms↗

Mulan: multiple-sequence local alignment and visualization for studying function and evolution.

Multiple-sequence alignment analysis is a powerful approach for understanding phylogenetic relationships, annotating genes, and detecting functional regulatory elements. With a growing number of partly or fully sequenced vertebrate genomes, effective tools for performing multiple comparisons are required to accurately and efficiently assist biological discoveries. Here we introduce Mulan (http://mulan.dcode.org/), a novel method and a network server for comparing multiple draft and finished-quality sequences to identify functional elements conserved over evolutionary time. Mulan brings together several novel algorithms: the TBA multi-aligner program for rapid identification of local sequence conservation, and the multiTF program for detecting evolutionarily conserved transcription factor binding sites in multiple alignments. In addition, Mulan supports two-way communication with the GALA database; alignments of multiple species dynamically generated in GALA can be viewed in Mulan, and conserved transcription factor binding sites identified with Mulan/multiTF can be integrated and overlaid with extensive genome annotation data using GALA. Local multiple alignments computed by Mulan ensure reliable representation of short- and large-scale genomic rearrangements in distant organisms. Mulan allows for interactive modification of critical conservation parameters to differentially predict conserved regions in comparisons of both closely and distantly related species. We illustrate the uses and applications of the Mulan tool through multispecies comparisons of the GATA3 gene locus and the identification of elements that are conserved in a different way in avians than in other genomes, allowing speculation on the evolution of birds. Source code for the aligners and the aligner-evaluation software can be freely downloaded from http://www.bx.psu.edu/miller_lab/.

Animals↗

Gene expression changes triggered by exposure of Haemophilus influenzae to novobiocin or ciprofloxacin: combined transcription and translation analysis.

The responses of Haemophilus influenzae to DNA gyrase inhibitors were analyzed at the transcriptional and the translational level. High-density microarrays based on the genomic sequence were used to monitor the expression levels of >80% of the genes in this bacterium. In parallel the proteins were analyzed by two-dimensional electrophoresis. DNA gyrase inhibitors of two different functional classes were used. Novobiocin, as a representative of one class, inhibits the ATPase activity of the enzyme, thereby indirectly changing the degree of DNA supercoiling. Ciprofloxacin, a representative of the second class, obstructs supercoiling by inhibiting the DNA cleavage-resealing reaction. Our results clearly show that different responses can be observed. Treatment with the ATPase inhibitor Novobiocin changed the expression rates of many genes, reflecting the fact that the initiation of transcription for many genes is sensitive to DNA supercoiling. Ciprofloxacin mainly stimulated the expression of DNA repair systems as a response to the DNA damage caused by the stable ternary complexes. In addition, changed expression levels were also observed for some genes coding for proteins either annotated as "unknown function" or "hypothetical" or for proteins not directly involved in DNA topology or repair.

Bacterial Proteins↗

Systematic identification of pseudogenes through whole genome expression evidence profiling.

The identification of pseudogenes is an integral and significant part of the genome annotation because of their abundance and their impact on the experimental analysis of functional genes. Most of the computational annotation systems are not optimized for systematic pseudogene recognition, often annotating pseudogenes as functional genes, and users then propagate these errors to subsequent analyses and interpretations. In order to validate gene annotations and to identify pseudogenes that are potentially mis-annotated, we developed a novel approach based on whole genome profiling of existing transcript and protein sequences. This method has two important features: (i) equally detects both processed and non-processed pseudogenes and (ii) can identify transcribed pseudogenes. Applying this method to the human Ensembl gene predictions, we discovered that 2011 (9% of total) Ensembl genes in the categories of known and novel might be pseudogenes based on expression evidence. Of these, 1200 genes are found to have no existing evidence of transcription, and 811 genes are found with transcription evidence but contain significant translation disruption. Approximately 40% of the 2011 identified pseudogenes presented a multi-exon structure, representing non-processed pseudogenes. We have demonstrated the power of whole genome profiling of expression sequences to improve the accuracy of gene annotations.

Computational Biology↗

PA-GOSUB: a searchable database of model organism protein sequences with their predicted Gene Ontology molecular function and subcellular localization.

PA-GOSUB (Proteome Analyst: Gene Ontology Molecular Function and Subcellular Localization) is a publicly available, web-based, searchable and downloadable database that contains the sequences, predicted GO molecular functions and predicted subcellular localizations of more than 107,000 proteins from 10 model organisms (and growing), covering the major kingdoms and phyla for which annotated proteomes exist (http://www.cs.ualberta.ca/~bioinfo/PA/GOSUB). The PA-GOSUB database effectively expands the coverage of subcellular localization and GO function annotations by a significant factor (already over five for subcellular localization, compared with Swiss-Prot v42.7), and more model organisms are being added to PA-GOSUB as their sequenced proteomes become available. PA-GOSUB can be used in three main ways. First, a researcher can browse the pre-computed PA-GOSUB annotations on a per-organism and per-protein basis using annotation-based and text-based filters. Second, a user can perform BLAST searches against the PA-GOSUB database and use the annotations from the homologs as simple predictors for the new sequences. Third, the whole of PA-GOSUB can be downloaded in either FASTA or comma-separated values (CSV) formats.

Amino Acid Sequence↗

metaExpertPro: A Computational Workflow for Metaproteomics Spectral Library Construction and Data-Independent Acquisition Mass Spectrometry Data Analysis.

Analysis of large-scale data-independent acquisition mass spectrometry metaproteomics data remains a computational challenge. Here, we present a computational pipeline called metaExpertPro for metaproteomics data analysis. This pipeline encompasses spectral library generation using data-dependent acquisition MS, protein identification and quantification using data-independent acquisition mass spectrometry, functional and taxonomic annotation, as well as quantitative matrix generation for both microbiota and hosts. By integrating FragPipe and DIA-NN, metaExpertPro offers compatibility with both Orbitrap and timsTOF MS instruments. To evaluate the depth and accuracy of identification and quantification, we conducted extensive assessments using human fecal samples and benchmark tests. Performance tests conducted on human fecal samples indicated that metaExpertPro quantified an average of 45,000 peptides in a 60-min diaPASEF injection. Notably, metaExpertPro outperformed three existing software tools by characterizing a higher number of peptides and proteins. Importantly, metaExpertPro maintained a low factual false discovery rate of approximately 5% for protein groups across four benchmark tests. Applying a filter of five peptides per genus, metaExpertPro achieved relatively high accuracy (F-score = 0.67-0.90) in genus diversity and showed a high correlation (rSpearman = 0.73-0.82) between the measured and true genus relative abundance in benchmark tests. Additionally, the quantitative results at the protein, taxonomy, and function levels exhibited high reproducibility and consistency across the commonly adopted public human gut microbial protein databases IGC and UHGP. In a metaproteomic analysis of dyslipidemia patients, metaExpertPro revealed characteristic alterations in microbial functions and potential interactions between the microbiota and the host.

Proteomics↗

NAVIP: Unraveling the influence of neighboring small sequence variants on functional impact prediction.

Once a suitable reference sequence has been generated, intra-species variation is often assessed by re-sequencing. Variant calling processes can reveal all differences between strains, accessions, genotypes, or individuals. These variants can be enriched with predictions about their functional implications based on available structural annotations, i.e., gene models. Although these functional impact predictions on a per-variant basis are often accurate, some challenging cases require the simultaneous incorporation of multiple adjacent variants into this prediction process. Examples include neighboring variants which modify each other's functional impact. The Neighborhood-Aware Variant Impact Predictor (NAVIP) considers all variants within a given protein coding sequence when predicting the effect. As a proof of concept, variants between the Arabidopsis thaliana accessions Columbia-0 and Niederzenz-1 were annotated. NAVIP is freely available on GitHub (https://github.com/bpucker/NAVIP) and accessible through a web server (https://pbb-tools.de).

Arabidopsis↗

Beyond Exons: Linking Noncoding Heritability and Polygenicity across Complex Human Traits and Disorders.

The genetic architecture of complex traits spans a continuum of polygenicity, yet it remains unclear how differences in polygenicity relate to the functional localization of SNP heritability across the genome. We use a MiXeR-based framework to partition heritability across exonic, intronic, and intergenic regions for 34 traits and introduce a likelihood-based annotation contribution score that quantifies annotation-specific impact on heritability. Exons explain a minority of heritability, and their contribution decreases with increasing polygenicity, from an average of 22% in less polygenic somatic diseases and biomarkers to 13% in highly polygenic psychiatric and cognitive phenotypes. Intergenic fractions show the opposite trend, whereas intronic fractions remain relatively stable. Analysis of a broader set of functional annotations reveals systematic differences along the polygenicity axis: highly polygenic traits show stronger contributions from comparative genomics and variant-effect scores, whereas less polygenic traits show stronger contributions in promoter, transcription, and chromatin annotations. Together, these results indicate that the functional partitioning of heritability systematically varies with polygenicity, pointing to a shift from gene-proximal regulatory architectures to architectures shaped by numerous dispersed regulatory effects as a key determinant of differences in polygenicity across traits.

Journal Article↗

Approaches to defining the ancestral eukaryotic protein complexome.

In this paper, we integrate and summarize the currently available information on the ancestral eukaryotic protein complexome, which is defined as the set of protein complexes that extant eukaryotes inherited from their last common ancestor. From the literature, we compiled lists of complexes with three or more distinct protein components from well-studied eukaryotic model organisms. Combinatorial complexes of membrane-associated signalling proteins and specific transcription factors were disregarded. A stringent but sensitive novel orthology detection algorithm, complemented with manual sequence similarity searches and with published data on whole genome or segmental and tandem gene duplications, enabled us to map the vast majority of these complexes to a virtual primitive eukaryote termed Eukaryotic Virtual Ancestor (EVA). EVA is intended to resemble the last common eukaryotic ancestor and to emulate the biological common denominator of the major extent eukaryotic lineages at the molecular level. The dataset was then used for the functional and domain annotation of the ancestral eukaryotic complexome. Furthermore, we illustrate its usefulness for inferring complexes of poorly studied eukaryotes and for the recognition of highly divergent orthologs. We also discuss the evolution of the circa 1,400 complex-associated ancestral proteins. As about 90% of these proteins have been conserved in all thirteen studied free-living eukaryotes, the evolutionary reduction and loss of complexes seems minimal. Moreover, the available data suggest that, in general, the acquisition of stable complexes of novel design occurs too slowly to be a major contributor to evolutionary innovation. Finally, given the stability of the ancestral eukarotic complexome we propose its use in the formulation of the mathematical systems that aim to simulate biological processes. Our data suggest that these simplified formulations can apply to most free-living model eukaryotes.

Animals↗

Overlapping deletions spanning the proximal two-thirds of the mouse t complex.

Chromosome deletion complexes in model organisms serve as valuable genetic tools for the functional and physical annotation of complex genomes. Among their many roles, deletions can serve as mapping tools for simple or quantitative trait loci (QTLs), genetic reagents for regional mutagenesis experiments, and, in the case of mice, models of human contiguous gene deletion syndromes. Deletions also are uniquely suited for identifying regions of the genome containing haploinsufficient or imprinted loci. Here we describe the creation of new deletions at the proximal end of mouse Chromosome (Chr) 17 by using the technique of ES cell irradiation and the extensive molecular characterization of these and previously isolated deletions that, in total, cover much of the mouse t complex. The deletions are arranged in five overlapping complexes that collectively span about 25 Mbp. Furthermore, we have integrated each of the deletion complexes with physical data from public and private mouse genome sequences, and our own genetic data, to resolve some discrepancies. These deletions will be useful for characterizing several phenomena related to the t complex and t haplotypes, including transmission ratio distortion, male infertility, and the collection of t haplotype embryonic lethal mutations. The deletions will also be useful for mapping other loci of interest on proximal Chr 17, including T-associated sex reversal ( Tas) and head-tilt ( het). The new deletions have thus far been used to localize the recently identified t haplolethal ( Thl1) locus to an approximately 1.3-Mbp interval.

Animals↗

Identification of immune-relevant genes from atlantic salmon using suppression subtractive hybridization.

In order to probe the interaction between an invading microorganism and its host, we have investigated differential gene expression in Atlantic salmon (Salmo salar) experimentally infected with the pathogen Aeromonas salmonicida, the causative agent of furunculosis. Subtractive cDNA libraries were constructed by suppression subtractive hybridization (SSH) from 3 immune-relevant tissues at 2 time points during the infection process. Both forward- and reverse-subtracted libraries were generated, and approximately 200 clones were sequenced from each library, giving a total of 1778 expressed sequence tags (ESTs), which were annotated according to functional categories and deposited in GenBank (BQ035314-BQ037059). Numerous genes involved in signal transduction, innate immunity, and other processes have been uncovered in the subtractive libraries. These include known acute-phase reactants, along with more novel genes encoding proteins such as tachylectin, hepcidin, precerebellin-like protein, O-methyltransferase, a putative saxitoxin-binding protein, and others. A subset of genes that were represented in the subtracted libraries was further analyzed by virtual Northern, or reverse transcription-polymerase chain reaction (RT-PCR) assays to verify their differential expression as a result of infection.

Aeromonas salmonicida↗

Identification of up-regulated genes after complete spinal cord transection in adult rats.

Spinal cord injury (SCI) initiates a cascade of events and these responses to injury are likely to be mediated and reflected by changes in mRNA concentrations. As a step towards understanding the complex mechanisms underlying repair and regeneration after SCI, the gene expression pattern was examined 4.5 days after complete transection at T8-9 level of rat spinal cord. Improved subtractive hybridization was used to establish a subtracted cDNA library using cDNAs from normal rat spinal cord as driver and cDNAs from injured spinal cord as tester. By expressed sequence tag (EST) sequencing, we obtained 73 EST fragments from this library, representing 40 differentially expressed genes. Among them, 32 were known genes and 8 were novel genes. Functions of all annotated genes were scattered in almost every important field of cell life such as DNA repair, detoxification, mRNA quality control, cell cycle control, and signaling, which reflected the complexity of SCI and regeneration. Then we verified subtraction results with semiquantitative RT-PCR for eight genes. These analyses confirmed, to a large extent, that the subtraction results accurately reflected the molecular changes occurring at 4.5 days post-SCI. The current study identified a number of genes that may shed new light on SCI-related inflammation, neuroprotection, neurite-outgrowth, synaptogenesis, and astrogliosis. In conclusion, the identification of molecular changes using improved subtractive hybridization may lead to a better understanding of molecular mechanisms responsible for repair and regeneration after SCI.

Animals↗

Comparative serial analysis of gene expression of transcript profiles of tomato roots infected with cyst nematode.

We analyzed global transcripts for tomato roots infected with the cyst nematode Globodera rostochiensis using serial analysis of gene expression (SAGE). SAGE libraries were made from nematode-infected roots and uninfected roots at 14 days after inoculation, and the clones including SAGE tags were sequenced. Genes were identified by matching the SAGE tags to tomato expressed sequence tags and cDNA databases. We then compiled a list of numerous genes according to the mRNA levels that were altered after cyst nematode infection. Our SAGE results showed significant changes in expression of many unreported genes involved in nematode infection. Of these, for discussion we selected five SAGE tags of RSI-1, BURP domain-containing protein, hexose transporter, P-rich protein, and PHAP2A that were activated by cyst nematode infection. Over 20% of the tags that were upregulated in the infected root have unknown functions (non-annotated), suggesting that we can obtain information on previously unreported and uncharacterized genes by SAGE. We can also obtain information on previously reported genes involved in nematode infection (e.g., multicystatin, peroxidase, catalase, pectin esterase, and S-adenosylmethionine transferase). To evaluate the validity of our SAGE results, seven genes were further analyzed by semiquantitative reverse transcriptase-polymerase chain reaction and Northern blot hybridization; the results agreed well with the SAGE data.

Animals↗

Investigating protein domain combinations in complete proteomes.

Protein-related information is more accumulated rather than reduced to a synthetic view. Itemising properties of protein sequences is informative, so is the list of ingredients to do some cooking, but without a recipe, that is, quantification and chronology, understanding is incomplete. If the goal of accumulating information is to discover or reveal the function and related biochemical mechanisms, information has to be weighed and ordered. As a guideline, the weight of a piece of information should reflect how often it consistently occurs in various contexts. We propose a common sense approach to quantify and put data and information into perspective. Complete bacterial proteomes are individually mapped with the Pfam-A database of domains and protein family signatures in an attempt to assess the modularity of proteins at the level of a single proteome and the implications of a modular description of proteins for a functional interpretation. Poorly annotated proteins in the most documented bacteria (E. coli and B. subtilis) were considered in an attempt to formulate hypothesis on the basis of domain/module content.

Bacillus subtilis↗

An atlas of differential gene expression during early Xenopus embryogenesis.

We have carried out a large-scale, semi-automated whole-mount in situ hybridization screen of 8369 cDNA clones in Xenopus laevis embryos. We confirm that differential gene expression is prevalent during embryogenesis since 24% of the clones are expressed non-ubiquitously and 8% are organ or cell type specific marker genes. Sequence analysis and clustering yielded 723 unique genes displaying a differential expression pattern. Of these, 18% were already described in Xenopus, 47% have homologs and 35% are lacking significant sequence similarity in databases. Many of them encode known developmental regulators. We classified 363 of the 723 genes for which a Gene Ontology annotation for molecular function could be attributed and found 'DNA binding' and 'enzyme' the most represented terms. The most common protein domains encoded in these embryonic, differentially expressed genes are the homeobox and RNA Recognition Motif (RRM). Fifty-nine putative orthologs of human disease genes, and 254 organ or cell specific marker genes were identified. Markers were found for nasal placode and archenteron roof, organs for which a specific marker was previously unavailable. Markers were also found for novel subdomains of various other organs. The tissues for which most markers were found are muscle and epidermis. Expression of cell cycle regulators fell in two classes, containing proliferation-promoting and anti-proliferative genes, respectively. We identified 66 new members of the BMP4, chromatin, endoplasmic reticulum, and karyopherin synexpression groups, thus providing a first glimpse of their probable cellular roles. Cluster analysis of tissues to measure tissue relatedness yielded some unorthodox affinities besides expectable lineage relationships. In conclusion, this study represents an atlas of gene expression patterns, which reveals embryonic regionalization, provides novel marker genes, and makes predictions about the functional role of unknown genes.

Animals↗

Use of SAGE technology to reveal changes in gene expression in Arabidopsis leaves undergoing cold stress.

The genes expressed within an organism determine its biological characteristics. Various internal or external factors can modulate these gene expression patterns, which then elicit physiological or pathological changes. We have characterized the global gene expression patterns of Arabidopsis leaves by serial analysis of gene expression (SAGE). A total of 21,280 SAGE tags were sequenced and 12,049 unique tags were identified. Among these, only 3367 tags (27.9%) were matched to the Arabidopsis cDNA or EST database. Functional analysis of annotated tags indicated that a significant proportion of the genes expressed in normal leaves were involved in energy and metabolism, especially in photosynthesis. To systematically analyze differential gene expression profiles under cold stress, a similar SAGE tag library from cold-treated leaves was constructed and analyzed. A comparison of the tags derived from the cold-treated leaves with those identified in the normal leaves revealed 272 differentially expressed genes (P<0.01): 82 genes were highly expressed in the normal leaves and 190 genes were highly expressed in the cold-treated leaves. After cold stress, in general, many of the genes involved in cell rescue/defense/cell death/aging, protein synthesis, metabolism, transport facilitation, and protein destination were induced. They included various COR genes, lipid transfer protein genes, alcohol dehydrogenase, beta-amylase and many novel genes. By comparison, down-regulated genes were mostly photosynthesis related genes involved in energy metabolism. The expression patterns of several cold responsive transcripts identified by SAGE were confirmed by northern analysis. The results presented here will provide valuable information for understanding the mechanisms of the freezing tolerance of plants.

Arabidopsis↗

RNAi living-cell microarrays for loss-of-function screens in Drosophila melanogaster cells.

RNA interference (RNAi)-mediated loss-of-function screening in Drosophila melanogaster tissue culture cells is a powerful method for identifying the genes underlying cell biological functions and for annotating the fly genome. Here we describe the development of living-cell microarrays for screening large collections of RNAi-inducing double-stranded RNAs (dsRNAs) in Drosophila cells. The features of the microarrays consist of clusters of cells 200 mum in diameter, each with an RNAi-mediated depletion of a specific gene product. Because of the small size of the features, thousands of distinct dsRNAs can be screened on a single chip. The microarrays are suitable for quantitative and high-content cellular phenotyping and, in combination screens, for the identification of genetic suppressors, enhancers and synthetic lethal interactions. We used a prototype cell microarray with 384 different dsRNAs to identify previously unknown genes that affect cell proliferation and morphology, and, in a combination screen, that regulate dAkt/dPKB phosphorylation in the absence of dPTEN expression.

Animals↗

An integrated data analysis approach to characterize genes highly expressed in hepatocellular carcinoma.

Hepatocellular carcinoma (HCC) is one of the major causes of cancer deaths worldwide. New diagnostic and therapeutic options are needed for more effective and early detection and treatment of this malignancy. We identified 703 genes that are highly expressed in HCC using DNA microarrays, and further characterized them in order to uncover novel tumor markers, oncogenes, and therapeutic targets for HCC. Using Gene Ontology annotations, genes with functions related to cell proliferation and cell cycle, chromatin, repair, and transcription were found to be significantly enriched in this list of highly expressed genes. We also identified a set of genes that encode secreted (e.g. GPC3, LCN2, and DKK1) or membrane-bound proteins (e.g. GPC3, IGSF1, and PSK-1), which may be attractive candidates for the diagnosis of HCC. A significant enrichment of genes highly expressed in HCC was found on chromosomes 1q, 6p, 8q, and 20q, and we also identified chromosomal clusters of genes highly expressed in HCC. The microarray analyses were validated by RT-PCR and PCR. This approach of integrating other biological information with gene expression in the analysis helps select aberrantly expressed genes in HCC that may be further studied for their diagnostic or therapeutic utility.

Biomarkers, Tumor↗