Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Widespread distribution of antisense transcripts in the Plasmodium falciparum genome.

The availability of the complete genome sequence of Plasmodium falciparum has facilitated high-throughput profiling of its complex life cycle, following the application of micro-array, proteomic, and serial analysis of gene expression (SAGE) technologies in this system. These, in turn, have yielded unprecedented insight into global gene expression, including the foremost demonstration of antisense transcription in the parasite. For example, owing to its inherent ability to sample novel ORFs and to predict transcript orientation, SAGE analysis in asexual forms led to the initial discovery of highly abundant antisense RNAs. To determine the extent of this phenomenon in P. falciparum, we have surveyed the distribution of both sense and antisense transcripts across the asexual transcriptome for the first time. To this end, a relational database integrating SAGE expression data with genome annotation information was constructed. This allowed the comprehensive annotation of a total of 17245 SAGE tags, extending over a 350-fold expression range. Transcripts from approximately 30% of the estimated 3D7 gene loci were present at detectable levels in mixed asexual stages, where loci involved in invasion and immune evasion; and carbohydrate metabolism were highly represented in the sense transcriptome. Approximately 12% of SAGE tags, however, were derived from the non-coding strand of nuclear-encoded ORFs, indicating that endogenous antisense RNAs are widespread in this system. Notably, these antisense transcripts were absent from the mitochondrial genome. Interestingly, we note that sense and antisense tag counts from single loci across the transcriptome were inversely related. Taken together, this data may provide first hints as to the possible function of antisense transcription in this system.

Animals↗

ProbeLynx: a tool for updating the association of microarray probes to genes.

As genome sequence data and gene prediction improve, probes developed for a given microarray experiment should be continuously re-evaluated for their specificity for given genes. ProbeLynx(www.pathogenomics.ca/probelynx) is a new web service which uses current genomic sequence information to re-examine microarray probe specificity and provide annotation updates relevant to determining which gene(s) and transcript(s) are associated with a given probe. Probe sequences (either oligonucleotide- or cDNA-based) are uploaded in FASTA format and the results returned as a tab-delimited flat file for insertion into a spreadsheet application or database management system for further analysis. ProbeLynx has been initially developed to focus on arrays derived from human, mouse, chicken and bovine genomes, but may be expanded to handle other genomic datasets. ProbeLynx offers microarray users the important ability to continuously assess the potential of a probe to cross-hybridize to paralogous genes and the suitability of a given probe to investigate a transcript of interest. By also including the latest gene function annotation information in the output, ProbeLynx provides the critical first step in updating microarray data annotation.

Animals↗

The Arabidopsis Information Resource (TAIR): a model organism database providing a centralized, curated gateway to Arabidopsis biology, research materials and community.

Arabidopsis thaliana is the most widely-studied plant today. The concerted efforts of over 11 000 researchers and 4000 organizations around the world are generating a rich diversity and quantity of information and materials. This information is made available through a comprehensive on-line resource called the Arabidopsis Information Resource (TAIR) (http://arabidopsis.org), which is accessible via commonly used web browsers and can be searched and downloaded in a number of ways. In the last two years, efforts have been focused on increasing data content and diversity, functionally annotating genes and gene products with controlled vocabularies, and improving data retrieval, analysis and visualization tools. New information include sequence polymorphisms including alleles, germplasms and phenotypes, Gene Ontology annotations, gene families, protein information, metabolic pathways, gene expression data from microarray experiments and seed and DNA stocks. New data visualization and analysis tools include SeqViewer, which interactively displays the genome from the whole chromosome down to 10 kb of nucleotide sequence and AraCyc, a metabolic pathway database and map tool that allows overlaying expression data onto the pathway diagrams. Finally, we have recently incorporated seed and DNA stock information from the Arabidopsis Biological Resource Center (ABRC) and implemented a shopping-cart style on-line ordering system.

Arabidopsis↗

BAV-LLPS: a database of bacterial, archaea, and virus liquid-liquid phase separation proteins.

MOTIVATION: Liquid-liquid phase separation (LLPS) is a key process underlying the formation of biomolecular condensates, such as membrane-less organelles, that compartmentalize biochemical processes inside the cells. While LLPS has been extensively studied in eukaryotes, its role in bacteria, archaea, and viruses remains far less characterized. Recent studies in bacteria have revealed that LLPS-driven condensates play critical roles in RNA processing, stress response, and pathogenicity. Similarly, many viruses exploit LLPS to facilitate crucial steps in their infection cycles, including viral entry, genome replication, assembly, and host immune evasion. RESULTS: In this work, we introduce a hand-curated database of LLPS proteins from bacteria, archaea, and viruses (BAV-LLPS Database). This resource, extended through sequence similarity searches, comprises over 5000 proteins and integrates diverse data including biological annotations, sequence features, predicted disordered regions, LLPS per site probability, and AlphaFold2-based structural models. Additionally, our web server enables users to explore both the curated and homologous derived datasets, providing a platform to uncover evolutionary relationships and intrinsic and differential properties of LLPS proteins across various taxonomic groups. This work seeks to deepen our understanding of LLPS mechanisms beyond eukaryotic organisms, emphasizing their significance across diverse life forms. It also aims to foster the development of specialized predictive tools that will facilitate the exploration and characterization of LLPS processes in a wide array of living organisms, thereby contributing to advancements in both fundamental biological research and applied biomedical sciences. AVAILABILITY AND IMPLEMENTATION: BAV-LLPS DB is freely accessible at https://bav-llps-db.bioinformatica.org/. The data can be retrieved from the website. The source code of the database can be downloaded from https://bav-llps-db.bioinformatica.org/download.

Databases, Protein↗

The Catalytic Site Atlas: a resource of catalytic sites and residues identified in enzymes using structural data.

The Catalytic Site Atlas (CSA) provides catalytic residue annotation for enzymes in the Protein Data Bank. It is available online at http://www.ebi.ac.uk/thornton-srv/databases/CSA. The database consists of two types of annotated site: an original hand-annotated set containing information extracted from the primary literature, using defined criteria to assign catalytic residues, and an additional homologous set, containing annotations inferred by PSI-BLAST and sequence alignment to one of the original set. The CSA can be queried via Swiss-Prot identifier and EC number, as well as by PDB code. CSA Version 1.0 contains 177 original hand- annotated entries and 2608 homologous entries, and covers approximately 30% of all EC numbers found in PDB. The CSA will be updated on a monthly basis to include homologous sites found in new PDBs, and new hand-annotated enzymes as and when their annotation is completed.

Animals↗

Alternate transcription of the Toll-like receptor signaling cascade.

BACKGROUND: Alternate splicing of key signaling molecules in the Toll-like receptor (Tlr) cascade has been shown to dramatically alter the signaling capacity of inflammatory cells, but it is not known how common this mechanism is. We provide transcriptional evidence of widespread alternate splicing in the Toll-like receptor signaling pathway, derived from a systematic analysis of the FANTOM3 mouse data set. Functional annotation of variant proteins was assessed in light of inflammatory signaling in mouse primary macrophages, and the expression of each variant transcript was assessed by splicing arrays. RESULTS: A total of 256 variant transcripts were identified, including novel variants of Tlr4, Ticam1, Tollip, Rac1, Irak1, 2 and 4, Mapk14/p38, Atf2 and Stat1. The expression of variant transcripts was assessed using custom-designed splicing arrays. We functionally tested the expression of Tlr4 transcripts under a range of cytokine conditions via northern and quantitative real-time polymerase chain reaction. The effects of variant Mapk14/p38 protein expression on macrophage survival were demonstrated. CONCLUSION: Members of the Toll-like receptor signaling pathway are highly alternatively spliced, producing a large number of novel proteins with the potential to functionally alter inflammatory outcomes. These variants are expressed in primary mouse macrophages in response to inflammatory mediators such as interferon-gamma and lipopolysaccharide. Our data suggest a surprisingly common role for variant proteins in diversification/repression of inflammatory signaling.

Alternative Splicing↗

Decoding the fine-scale structure of a breast cancer genome and transcriptome.

A comprehensive understanding of cancer is predicated upon knowledge of the structure of malignant genomes underlying its many variant forms and the molecular mechanisms giving rise to them. It is well established that solid tumor genomes accumulate a large number of genome rearrangements during tumorigenesis. End Sequence Profiling (ESP) maps and clones genome breakpoints associated with all types of genome rearrangements elucidating the structural organization of tumor genomes. Here we extend the ESP methodology in several directions using the breast cancer cell line MCF-7. First, targeted ESP is applied to multiple amplified loci, revealing a complex process of rearrangement and co-amplification in these regions reminiscent of breakage/fusion/bridge cycles. Second, genome breakpoints identified by ESP are confirmed using a combination of DNA sequencing and PCR. Third, in vitro functional studies assign biological function to a rearranged tumor BAC clone, demonstrating that it encodes anti-apoptotic activity. Finally, ESP is extended to the transcriptome identifying four novel fusion transcripts and providing evidence that expression of fusion genes may be common in tumors. These results demonstrate the distinct advantages of ESP including: (1) the ability to detect all types of rearrangements and copy number changes; (2) straightforward integration of ESP data with the annotated genome sequence; (3) immortalization of the genome; (4) ability to generate tumor-specific reagents for in vitro and in vivo functional studies. Given these properties, ESP could play an important role in a tumor genome project.

Breast Neoplasms↗

dbscATAC: a resource of single-cell super-enhancers/enhancers and gene markers derived from scATAC-seq data.

MOTIVATION: scATAC-seq enables high-resolution mapping of cis-regulatory elements. It has been widely applied to uncover cell-type-specific regulatory networks and complement scRNA-seq analysis in numerous studies. However, a large number of datasets generated by scATAC-seq remain underutilized due to limited exploration of super-enhancers/typical enhancers and gene markers. A comprehensive resource enabling cell-type-specific annotation of cis-regulatory elements and their dynamic enhancer-gene linkages remains an urgent unmet need for scATAC-seq. RESULTS: We present dbscATAC, a specialized single-cell database for annotating super-enhancers, gene markers, and enhancer-gene interactions derived from scATAC-seq data. Using improved machine learning algorithms, we identified 213 835 super-enhancers across 520 tissue/cell types from three species, as well as 347 484 gene markers, 13 470 526 enhancers, and 10 402 346 enhancer-gene interactions derived from 1 668 076 single cells spanning 1028 tissue/cell types in 13 species. An easy-to-use online platform with multiple analytic modules and hierarchical query options was developed for searching, browsing and visualizing single-cell super-enhancers, enhancers, and gene markers. dbscATAC provides a comprehensive resource to facilitate the exploration of enhancer landscapes, gene regulation, and cell-type-specific characteristics in single-cell epigenomics. AVAILABILITY AND IMPLEMENTATION: The database with all the super-enhancer/enhancer annotation data is available at http://singlecelldb.com/dbscATAC/index.php. And the source code of dbscATAC for prediction of SEs, enhancers, and gene markers are available at https://github.com/EvansGao/dbscATAC. The source code, tissue/cell type description, and data summary can be downloaded at DOI: 10.6084/m9.figshare.28706414.scATAC-seq, Database, Super-enhancers/enhancers, Gene markers.

Enhancer Elements, Genetic↗

T1DBase, a community web-based resource for type 1 diabetes research.

T1DBase (http://T1DBase.org) is a public website and database that supports the type 1 diabetes (T1D) research community. The site is currently focused on the molecular genetics and biology of T1D susceptibility and pathogenesis. It includes the following datasets: annotated genome sequence for human, rat and mouse; information on genetically identified T1D susceptibility regions in human, rat and mouse, and genetic linkage and association studies pertaining to T1D; descriptions of NOD mouse congenic strains; the Beta Cell Gene Expression Bank, which reports expression levels of genes in beta cells under various conditions, and annotations of gene function in beta cells; data on gene expression in a variety of tissues and organs; and biological pathways from KEGG and BioCarta. Tools on the site include the GBrowse genome browser, site-wide context dependent search, Connect-the-Dots for connecting gene and other identifiers from multiple data sources, Cytoscape for visualizing and analyzing biological networks, and the GESTALT workbench for genome annotation. All data are open access and all software is open source.

Animals↗

Characterizing the metabolic phenotype: a phenotype phase plane analysis.

Genome-scale metabolic maps can be reconstructed from annotated genome sequence data, biochemical literature, bioinformatic analysis, and strain-specific information. Flux-balance analysis has been useful for qualitative and quantitative analysis of metabolic reconstructions. In the past, FBA has typically been performed in one growth condition at a time, thus giving a limited view of the metabolic capabilities of a metabolic network. We have broadened the use of FBA to map the optimal metabolic flux distribution onto a single plane, which is defined by the availability of two key substrates. A finite number of qualitatively distinct patterns of metabolic pathway utilization were identified in this plane, dividing it into discrete phases. The characteristics of these distinct phases are interpreted using ratios of shadow prices in the form of isoclines. The isoclines can be used to classify the state of the metabolic network. This methodology gives rise to a "phase plane" analysis of the metabolic genotype-phenotype relation relevant for a range of growth conditions. Phenotype phase planes (PhPPs) were generated for Escherichia coli growth on two carbon sources (acetate and glucose) at all levels of oxygenation, and the resulting optimal metabolic phenotypes were studied. Supplementary information can be downloaded from our website (http://epicurus.che.udel.edu).

Computational Biology↗

Statistical analysis and prediction of protein-protein interfaces.

Predicting protein-protein interfaces from a three-dimensional structure is a key task of computational structural proteomics. In contrast to geometrically distinct small molecule binding sites, protein-protein interface are notoriously difficult to predict. We generated a large nonredundant data set of 1494 true protein-protein interfaces using biological symmetry annotation where necessary. The data set was carefully analyzed and a Support Vector Machine was trained on a combination of a new robust evolutionary conservation signal with the local surface properties to predict protein-protein interfaces. Fivefold cross validation verifies the high sensitivity and selectivity of the model. As much as 97% of the predicted patches had an overlap with the true interface patch while only 22% of the surface residues were included in an average predicted patch. The model allowed the identification of potential new interfaces and the correction of mislabeled oligomeric states.

Animals↗

A genomics approach to crop pest and disease research.

Genome-wide analyses of gene function and gene expression are beginning to yield valuable information in many areas of biological research, and these genomic tools are now being applied to crop pest and disease research. DNA sequencing of cDNA libraries to generate sets of expressed sequence tags (ESTs) are allowing gene compendiums for crop diseases to be compiled. Annotation of such data collections is also providing a wealth of functional information about gene products through similarities to proteins with known function. The next phase of the functional genomics era will be to employ large-scale techniques to knock out or silence genes in order to synthesize gene-specific mutants for phenotypic analysis and to use micro-array methodology to analyze global gene expression, protein turnover and protein processing during the processes of parasitism and colonization. Application of these technologies promises to accelerate the pace that biological information relevant to crop protection accrues. The ability of researchers to assimilate this information into complex models and workable hypotheses is, thus, set to revolutionize the way we study pests and diseases of crop plants.

Crops, Agricultural↗

Characterization of the human Xq21.3/Yp11 homology block and conservation of organization in primates.

The Xq21.3/Yp11 homology block on the human sex chromosomes represents a recent addition to the Y chromosome through a transposition event. It is believed that this transfer of material occurred after the divergence of the hominid lineage from other great apes. In this paper we investigate the structure and evolution of the block through fluorescence in situ hybridisation, contig assembly, the polymerase chain reaction, exon trapping, sequence comparison, and annotation of sequence data. The overall structure is well conserved between the human X chromosome and the Y chromosome as well as between the X chromosomes from different primates. Although the sequence data reveal a high level of nucleotide sequence identity for the human X and Y, there are regions of significant divergence, such as that around the marker DXS214. These are presumably the consequence of multiple rearrangements during evolution and are of particular importance with respect to the potential gene content in this segment of the interval.

Animals↗

Cytosine methylation is not the major factor inducing CpG dinucleotide deficiency in bacterial genomes.

CpG dinucleotide deficiency has been found in viruses, mitochondria, prokaryotes, and eukaryotes. The consensual explanation is that it is due to deamination of methylated cytosines, as established for vertebrate and plants. However, we still do not know whether C5 cytosine methylation is also the major cause of CpG deficiency in bacteria. By combining annotation and experimental data identifying the presence of C5 cytosine methyltransferases with analysis of CpG relative abundance in 67 bacterial species, we found that CpG relative abundance in most bacterial genomes that have cytosine C5 methyltransferases tends to be in the normal range (observed/expected values between 0.82 and 1.21). In contrast, many bacterial species likely to be lacking C5 cytosine methylation showed CpG deficiency. Furthermore, when comparing genomes with one another, TpG and CpA relative abundances were found to be independent from CpG relative abundance. This contrasted with intragenome analyses, where C3pG1 relative abundance (the subscripts refer to position of a nucleotide in a codon) was found to be generally positively correlated with T3pG1 relative abundances when plotted against GC content in protein coding sequences (CDSs). This suggests the existence of alternative mechanisms contributing to CpG deficiency in bacteria.

Bacteria↗

Computer database of ambulatory EEG signals.

The paper describes an ambulatory EEG database. The database contains segments of AEEGs done on 45 subjects. Each epoch (1/8th second or more) of AEEG data has been annotated into 1 of 40 classes. The classes represent background activity, paroxysmal patterns and artifacts. The majority of classes have over 200 discrete epochs. The structure is flexible enough to allow additional epochs to be readily added. The database is stored on transportable media such as digital magnetic tape or hard disk and is thus available to other researchers in the field. The database can be used to design, evaluate and compare EEG signal processing algorithms and pattern recognition systems. It can also serve as an educational medium in EEG laboratories.

Adolescent↗

SIGNATURE: a single-particle selection system for molecular electron microscopy.

SIGNATURE is a particle selection system for molecular electron microscopy. It applies a hierarchical screening procedure to identify molecular particles in EM micrographs. The user interface of the program provides versatile functions to facilitate image data visualization, particle annotation and particle quality inspection. The system design emphasizes both functionality and usability. This software has been released to the EM community and has been successfully applied to macromolecular structural analyses.

Algorithms↗

GXXXG and GXXXA motifs stabilize FAD and NAD(P)-binding Rossmann folds through C(alpha)-H... O hydrogen bonds and van der waals interactions.

Here we present evidence that domains in soluble proteins containing either the GXXXG or GXXXA motif are stabilized by the interaction of a beta-strand with the following alpha-helix. As an example, we characterized a beta-strand-helix interaction from the FAD or NAD(P)-binding Rossmann fold. The Rossmann fold is one of the three most highly represented folds in the Protein Data Bank (PDB). A subset of the proteins that adopt the Rossmann fold also bind to nucleotide cofactors such as FAD and NAD(P) and function as oxidoreductases. These Rossmann folds can often be identified by the short amino acid sequence motif, GX(1-2)GXXG. Here, we present evidence that in addition to this sequence motif, Rossmann folds that bind FAD and NAD(P) also typically contain either GXXXG or GXXXA motifs, where the first glycyl residue of these motifs and the third glycyl residue of the GX(1-2)GXXG motif are the same residue. These two motifs appear to stabilize the Rossmann fold: the first glycyl residue of either the GXXXG or GXXXA motif contacts the carbonyl oxygen atom from the first glycyl residue of the GX(1-2)GXXG motif consistent with the formation of a C(alpha)-H cdots, three dots, centered O hydrogen bond. In addition, both the glycyl and alanyl residues of the GXXXG or GXXXA motifs form van der Waals interactions with either a valine or isoleucine residue located either seven or eight residues further back along the polypeptide chain from the first glycine of the GXXXG or GXXXA motifs. Therefore, we combine both the GX(1-2)GXXG and GXXXG/A motifs into an extended motif, V/IXGX(1-2)GXXGXXXG/A, that is more strongly indicative than previously described motifs of Rossmann folds that bind FAD or NAD(P). The V/IXGX(1-2)GXXGXXXG/A motif can be used to search genomic sequence data and to annotate the function of proteins containing the motif as oxidoreductases, including proteins of previously unknown function.

Amino Acid Motifs↗

Adult mouse brain gene expression patterns bear an embryologic imprint.

The current model to explain the organization of the mammalian nervous system is based on studies of anatomy, embryology, and evolution. To further investigate the molecular organization of the adult mammalian brain, we have built a gene expression-based brain map. We measured gene expression patterns for 24 neural tissues covering the mouse central nervous system and found, surprisingly, that the adult brain bears a transcriptional "imprint" consistent with both embryological origins and classic evolutionary relationships. Embryonic cellular position along the anterior-posterior axis of the neural tube was shown to be closely associated with, and possibly a determinant of, the gene expression patterns in adult structures. We also observed a significant number of embryonic patterning and homeobox genes with region-specific expression in the adult nervous system. The relationships between global expression patterns for different anatomical regions and the nature of the observed region-specific genes suggest that the adult brain retains a degree of overall gene expression established during embryogenesis that is important for regional specificity and the functional relationships between regions in the adult. The complete collection of extensively annotated gene expression data along with data mining and visualization tools have been made available on a publicly accessible web site (www.barlow-lockhart-brainmapnimhgrant.org).

Algorithms↗