Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

ESTIMA, a tool for EST management in a multi-project environment.

BACKGROUND: Single-pass, partial sequencing of complementary DNA (cDNA) libraries generates thousands of chromatograms that are processed into high quality expressed sequence tags (ESTs), and then assembled into contigs representative of putative genes. Usually, to be of value, ESTs and contigs must be associated with meaningful annotations, and made available to end-users. RESULTS: A web application, Expressed Sequence Tag Information Management and Annotation (ESTIMA), has been created to meet the EST annotation and data management requirements of multiple high-throughput EST sequencing projects. It is anchored on individual ESTs and organized around different properties of ESTs including chromatograms, base-calling quality scores, structure of assembled transcripts, and multiple sources of comparison to infer functional annotation, Gene Ontology associations, and cDNA library information. ESTIMA consists of a relational database schema and a set of interactive query interfaces. These are integrated with a suite of web-based tools that allow a user to query and retrieve information. Further, query results are interconnected among the various EST properties. ESTIMA has several unique features. Users may run their own EST processing pipeline, search against preferred reference genomes, and use any clustering and assembly algorithm. The ESTIMA database schema is very flexible and accepts output from any EST processing and assembly pipeline. ESTIMA has been used for the management of EST projects of many species, including honeybee (Apis mellifera), cattle (Bos taurus), songbird (Taeniopygia guttata), corn rootworm (Diabrotica vergifera), catfish (Ictalurus punctatus, Ictalurus furcatus), and apple (Malus x domestica). The entire resource may be downloaded and used as is, or readily adapted to fit the unique needs of other cDNA sequencing projects. CONCLUSIONS: The scripts used to create the ESTIMA interface are freely available to academic users in an archived format from http://titan.biotec.uiuc.edu/ESTIMA/. The entity-relationship (E-R) diagrams and the programs used to generate the Oracle database tables are also available. We have also provided detailed installation instructions and a tutorial at the same website. Presently the chromatograms, EST databases and their annotations have been made available for cattle and honeybee brain EST projects. Non-academic users need to contact the W.M. Keck Center for Functional and Comparative Genomics, University of Illinois at Urbana-Champaign, Urbana, IL, for licensing information.

Animals↗

Generation of an improved cytogenetic and comparative map of Bos taurus chromosome BTA27.

Comparative genome analysis in cattle, human, and mouse identified various evolutionary breakpoints between Bos taurus 27 chromosome (BTA27) and corresponding segments in the Homo sapiens 4 and 8 chromosomes (HSA4, HSA8) and the Mus musculus 8 chromosome (MMU8). The fragmentary cytogenetic location of breaks is based on nine known loci and Zoo-FISH data on BTA27. A comparative mapping approach combining in-silico mapping and physical mapping by fluorescence in-situ hybridization (FISH) revealed an improved cytogenetic map of BTA27 based on 25 new and nine existing assignments of loci. Furthermore, hybrid cell mapping techniques identified and anchored three additional gene loci on BTA27. The BTA27 map was compared with available mapping and annotated sequence data for the chromosome and a generated comparative map displays conserved syntenic chromosome blocks between cattle, human, and mouse. The new anchor loci identify and narrow down evolutionary breakpoints on a cytogenetic level and can help to support the cattle genome assembly and annotation process.

Animals↗

The genexpress IMAGE knowledge base of the human muscle transcriptome: a resource of structural, functional, and positional candidate genes for muscle physiology and pathologies.

Sequence, gene mapping, and expression data corresponding to 910 genes transcribed in human skeletal muscle have been integrated to form the muscle module of the Genexpress IMAGE Knowledge Base. Based on cDNA array hybridization, a set of 14 transcripts preferentially or specifically expressed in muscle have been selected and characterized in more detail: Their pattern of expression was confirmed by Northern blot analysis; their structure was further characterized by full-insert cDNA sequencing and cDNA extension; the map location of the corresponding genes was refined by radiation hybrid mapping. Five of the 14 selected genes appear as interesting positional and functional candidate genes to study in relation with muscle physiology and/or specific orphan muscular pathologies. One example is discussed in more detail. The expression profiling data and the associated Genexpress Index2 entries for the 910 genes and the detailed characterization of the 14 selected transcripts are available from a dedicated Web server at. The database has been organized to provide the users with a working space where they can find curated, annotated, integrated data for their genes of interest. Different navigation routes to exploit the resource are discussed.

Base Sequence↗

ABI sequencing analysis. Manipulation of sequence data from the ABI DNA sequencer.

The ABI Sequencing Analysis application is designed specifically for the analysis of data produced by the ABI DNA Sequencer. The ABI sequencer is a laser-based instrument that utilizes fluorescent labels to analyze the products of a sequencing reaction as they migrate through a gel. After the data are collected from a sequencing run, the Analysis program identifies and tracks the sample lanes of the gel and subsequently normalizes and integrates the raw data into a chromatogram of the final sequence. For the user, there are basically two types of files that can be manipulated to potentially improve the analysis results. The Gel File consists of a computer generated image of the sequencing gel with the fluorescent DNA banding patterns. This image allows the user to view and edit the tracking lines generated and used by Analysis to collect data points for each sample. Individual Sample Files are stored for each of the samples analyzed and include the chromatogram, raw data, and annotations and information regarding the sample and sequence run. Generally, the products of a sequencing reaction are easily resolved and the Analysis software interprets the correct nucleotide sequence. Ambiguous base calls tend to occur near the end of the sequence and may be either edited or deleted by the user before exporting the data for further comparisons or alignments. Occasionally the tracking lines within the gel image may need to be adjusted or moved. The sample data are then reextracted from the Gel File and analyzed again. This review explains the general operation of Analysis in terms of viewing and editing a chromatogram, retracking the lanes of a Gel File, and analyzing the final sample data. The three versions 1.2.1, 2.1.2, and 3.3 are discussed.

Animals↗

Robust sensor fusion improves heart rate estimation: clinical evaluation.

OBJECTIVE: To determine if Robust Sensor Fusion (RSF), a method designed to fuse data from multiple sensors with redundant heart rate information can be used to improve the quality of heart rate data. To determine if the improved estimate of heart rate can reduce the number of false and missed heart rate alarms. METHODS: A total of 85 monitoring periods were investigated, 12 from the operating room, 60 from adult ICU and 13 from pediatric ICU. The operating room periods began with induction of anesthesia and ended at the completion of the anesthetic. For the ICU data, four hour blocks of time were studied. For each monitoring period, HR values were recorded at 5 second intervals or less from the ECG, SpO2 and IAC using a SpaceLabs Medical Gateway connected to a SpaceLabs Medical PC2. Fused estimates of HR were derived for every time point using RSF and all results accepted regardless of confidence value. Data were annotated manually to identify the "reference" HR (that HR value most likely to be correct) at all time points. All HR values from the sensors and the fused estimate that were different from the reference HR by more than +/- 5 beats/min were considered inaccurate. For each monitoring period, the total time per hour that data were either inaccurate or unavailable was calculated for each sensor as well as the fused estimates. The total time of false and missed HR alarms was found for all sensors and the fused estimate by comparing the data to thresholds for both high and low HR alarms at 150 bpm, 130 bpm, 110 bpm and 50 bpm, 40 bpm, 30 bpm respectively. RESULTS: The fused estimate of HR was consistently as good or better than the estimate available from any individual sensor. The fused estimates also consistently reduced the incidence of false alarms compared with individual sensors without an unacceptable incidence of missed alarms. DISCUSSION: Redundancy in sensor measurements can be used to improve HR estimation in the clinical setting. Methods like RSF which improve the quality of monitored data and reduce nuisance alarms will enhance the value of patient monitors to clinicians.

Adult↗

Extending traditional query-based integration approaches for functional characterization of post-genomic data.

MOTIVATION: To identify and characterize regions of functional interest in genomic sequence requires full, flexible query access to an integrated, up-to-date view of all related information, irrespective of where it is stored (within an organization or across the Internet) and its format (traditional database, flat file, web site, results of runtime analysis). Wide-ranging multi-source queries often return unmanageably large result sets, requiring non-traditional approaches to exclude extraneous data. RESULTS: Target Informatics Net (TINet) is a readily extensible data integration system developed at GlaxoSmith- Kline (GSK), based on the Object-Protocol Model (OPM) multidatabase middleware system of Gene Logic Inc. Data sources currently integrated include: the Mouse Genome Database (MGD) and Gene Expression Database (GXD), GenBank, SwissProt, PubMed, GeneCards, the results of runtime BLAST and PROSITE searches, and GSK proprietary relational databases. Special-purpose class methods used to filter and augment query results include regular expression pattern-matching over BLAST HSP alignments and retrieving partial sequences derived from primary structure annotations. All data sources and methods are accessible through an SQL-like query language or a GUI, so that when new investigations arise no additional programming beyond query specification is required. The power and flexibility of this approach are illustrated in such integrated queries as: (1) 'find homologs in genomic sequence to all novel genes cloned and reported in the scientific literature within the past three months that are linked to the MeSH term 'neoplasms"; (2) 'using a neuropeptide precursor query sequence, return only HSPs where the target genomic sequences conserve the G[KR][KR] motif at the appropriate points in the HSP alignment'; and (3) 'of the human genomic sequences annotated with exon boundaries in GenBank, return only those with valid putative donor/acceptor sites and start/stop codons'.

Animals↗

A novel deep learning-driven framework for improving lncRNA comprehensive annotation with LncADeep 2.0.

MOTIVATION: Long non-coding RNAs (lncRNAs) have emerged as crucial players in diverse physiological and pathological processes, yet the biological mechanisms of the vast majority of lncRNAs remain elusive. To fill this gap, it is necessary to improve the accuracy of lncRNA identification and functional annotation. RESULTS: Here, we introduce LncADeep 2.0, an integrated deep learning framework designed to meet these needs. In the identification module, LncADeep 2.0 incorporated novel peptide features along with sequence and structural information, demonstrating superior performance over our previous LncADeep and other existing tools on both annotated transcripts from GENCODE and RNA-seq data. For functional annotation, LncADeep 2.0 leveraged lncRNA-centric interaction networks and gene ontology terms through the transfer learning strategy to achieve robust annotation performance with limited functional data. Compared to LncADeep, LncADeep 2.0 could accurately elucidate the general functions of given lncRNA sequences, predict tissue- or cell-type-specific functions from bulk and single-cell RNA-seq data, and establish connections between tumor-associated lncRNAs and genomic markers. Overall, LncADeep 2.0 stands out as an efficient and reliable tool for lncRNA identification and functional annotation across a wide spectrum of biological processes. AVAILABILITY AND IMPLEMENTATION: LncADeep 2.0 is available for use at https://github.com/Jefferson-Chou/LncADeep2 and https://doi.org/10.5281/zenodo.17164767.

RNA, Long Noncoding↗

Biomedical data integration: using XML to link clinical and research data sets.

Data integration occurs when a query proceeds through multiple data sets, thereby relating diverse data extracted from different data sources. Data integration is particularly important to biomedical researchers since data obtained from experiments on human tissue specimens have little applied value unless they can be combined with medical data (i.e., pathologic and clinical information). In the past, research data were correlated with medical data by manually retrieving, reading, assembling and abstracting patient charts, pathology reports, radiology reports and the results of special tests and procedures. Manual annotation of research data is impractical when experiments involve hundreds or thousands of tissue specimens resulting in large, complex data collections. The purpose of this paper is to review how XML (eXtensible Markup Language) provides the fundamental tools that support biomedical data integration. The article also discusses some of the most important challenges that block the widespread availability of annotated biomedical data sets.

Data Collection↗

Annotating nonspecific SAGE tags with microarray data.

SAGE (serial analysis of gene expression) detects transcripts by extracting short tags from the transcripts. Because of the limited length, many SAGE tags are shared by transcripts from different genes. Relying on sequence information in the general gene expression database has limited power to solve this problem due to the highly heterogeneous nature of the deposited sequences. Considering that the complexity of gene expression at a single tissue level should be much simpler than that in the general expression database, we reasoned that by restricting gene expression to tissue level, the accuracy of gene annotation for the nonspecific SAGE tags should be significantly improved. To test the idea, we developed a tissue-specific SAGE annotation database based on microarray data (). This database contains microarray expression information represented as UniGene clusters for 73 normal human tissues and 18 cancer tissues and cell lines. The nonspecific SAGE tag is first matched to the database by the same tissue type used by both SAGE and microarray analysis; then the multiple UniGene clusters assigned to the nonspecific SAGE tag are searched in the database under the matched tissue type. The UniGene cluster presented solely or at higher expression levels in the database is annotated to represent the specific gene for the nonspecific SAGE tags. The accuracy of gene annotation by this database was largely confirmed by experimental data. Our study shows that microarray data provide a useful source for annotating the nonspecific SAGE tags.

Cell Line↗

The Diatom EST Database.

The Diatom EST database provides integrated access to expressed sequence tag (EST) data from two eukaryotic microalgae of the class Bacillariophyceae, Phaeodactylum tricornutum and Thalassiosira pseudonana. The database currently contains sequences of close to 30,000 ESTs organized into PtDB, the P.tricornutum EST database, and TpDB, the T.pseudonana EST database. The EST sequences were clustered and assembled into a non-redundant set for each organism, and these non-redundant sequences were then subjected to automated annotation using similarity searches against protein and domain databases. EST sequences, clusters of contiguous sequences, their annotation and analysis with reference to the publicly available databases, and a codon usage table derived from a subset of sequences from PtDB and TpDB can all be accessed in the Diatom EST Database. The underlying RDBMS enables queries over the raw and annotated EST data and retrieval of information through a user-friendly web interface, with options to perform keyword and BLAST searches. The EST data can also be retrieved based on Pfam domains, Cluster of Orthologous Groups (COG) and Gene Ontologies (GO) assigned to them by similarity searches. The Database is available at http://avesthagen.sznbowler.com.

DNA, Algal↗

Marine genomics: a clearing-house for genomic and transcriptomic data of marine organisms.

BACKGROUND: The Marine Genomics project is a functional genomics initiative developed to provide a pipeline for the curation of Expressed Sequence Tags (ESTs) and gene expression microarray data for marine organisms. It provides a unique clearing-house for marine specific EST and microarray data and is currently available at http://www.marinegenomics.org. DESCRIPTION: The Marine Genomics pipeline automates the processing, maintenance, storage and analysis of EST and microarray data for an increasing number of marine species. It currently contains 19 species databases (over 46,000 EST sequences) that are maintained by registered users from local and remote locations in Europe and South America in addition to the USA. A collection of analysis tools are implemented. These include a pipeline upload tool for EST FASTA file, sequence trace file and microarray data, an annotative text search, automated sequence trimming, sequence quality control (QA/QC) editing, sequence BLAST capabilities and a tool for interactive submission to GenBank. Another feature of this resource is the integration with a scientific computing analysis environment implemented by MATLAB. CONCLUSION: The conglomeration of multiple marine organisms with integrated analysis tools enables users to focus on the comprehensive descriptions of transcriptomic responses to typical marine stresses. This cross species data comparison and integration enables users to contain their research within a marine-oriented data management and analysis environment.

Animals↗

[Application of Excel Visual Basic for efficiently complete statistic analysis].

OBJECTIVE: In order to analyze multiple statistic tables more efficiently Excel Visual Basic for Application (VBA) was introduced through the use of an example of calculating standardized mortality rates (SMRs). METHODS: Mortality data of cancer and cardiovascular diseases, by sex and age, have been collected from 1991 to 2003 by the Center for Disease Control and Prevention of Shanghai Huangpu District. Standard population composition was defined as Chinese census statistics in 2000. The male's SMRs were calculated, using Excel VBA for each year and classification of cancers. RESULTS: The male's SMRs were obtained by year and different cancers. At the same time, the results were listed in the cancer's SMRs table for male. CONCLUSIONS: Excel is more flexible than general database on the combination of data and annotation. Excel VBA is better than the basic Excel in operating multiple tables simultaneously and man-machine conversation. Statistic analysis can be efficiently completed by using Excel VBA.

Age Factors↗

The Database of Quantitative Cellular Signaling: management and analysis of chemical kinetic models of signaling networks.

MOTIVATION: Analysis of cellular signaling interactions is expected to pose an enormous informatics challenge, perhaps even larger than analyzing the genome. The complex networks arising from signaling processes are traditionally represented as block diagrams. A key step in the evolution toward a more quantitative understanding of signaling is to explicitly specify the kinetics of all chemical reaction steps in a pathway. Technical advances in proteomics and high-throughput protein interaction assays promise a flood of such quantitative data. While annotations, molecular information and pathway connectivity have been compiled in several databases, and there are several proposals for general cell model description languages, there is currently little experience with databases of chemical kinetics and reaction level models of signaling networks. RESULTS: The Database of Quantitative Cellular Signaling is a repository of models of signaling pathways. It is intended both to serve the growing field of chemical-reaction level simulation of signaling networks, and to anticipate issues in large-scale data management for signaling chemistry. AVAILABILITY: The Database of Quantitative Cellular Signaling is available at http://doqcs.ncbs.res.in. Links to the signaling model simulator, GENESIS/Kinetikit are at http://www.ncbs.res.in/~bhalla/kkit/index.html and are also provided from within the database. The database source code is available under the GNU Public License.

Abstracting and Indexing↗

Query3d: a new method for high-throughput analysis of functional residues in protein structures.

BACKGROUND: The identification of local similarities between two protein structures can provide clues of a common function. Many different methods exist for searching for similar subsets of residues in proteins of known structure. However, the lack of functional and structural information on single residues, together with the low level of integration of this information in comparison methods, is a limitation that prevents these methods from being fully exploited in high-throughput analyses. RESULTS: Here we describe Query3d, a program that is both a structural DBMS (Database Management System) and a local comparison method. The method conserves a copy of all the residues of the Protein Data Bank annotated with a variety of functional and structural information. New annotations can be easily added from a variety of methods and known databases. The algorithm makes it possible to create complex queries based on the residues' function and then to compare only subsets of the selected residues. Functional information is also essential to speed up the comparison and the analysis of the results. CONCLUSION: With Query3d, users can easily obtain statistics on how many and which residues share certain properties in all proteins of known structure. At the same time, the method also finds their structural neighbours in the whole PDB. Programs and data can be accessed through the PdbFun web interface.

Algorithms↗

DDBJ in collaboration with mass-sequencing teams on annotation.

In the past year, we at DDBJ (DNA Data Bank of Japan; http://www.ddbj.nig.ac.jp) collected and released 1,066,084 entries or 718,072,425 bases including the whole chromosome 22 of chimpanzee, the whole-genome shotgun sequences of silkworm and various others. On the other hand, we hosted workshops for human full-length cDNA annotation and participated in jamborees of mouse full-length cDNA annotation. The annotated data are made public at DDBJ. We are also in collaboration with a RIKEN team to accept and release the CAGE (Cap Analysis Gene Expression) data under a new category, MGA (Mass Sequences for Genome Annotation). The data will be useful for studying gene expression control in many aspects.

Animals↗

XRate: a fast prototyping, training and annotation tool for phylo-grammars.

BACKGROUND: Recent years have seen the emergence of genome annotation methods based on the phylo-grammar, a probabilistic model combining continuous-time Markov chains and stochastic grammars. Previously, phylo-grammars have required considerable effort to implement, limiting their adoption by computational biologists. RESULTS: We have developed an open source software tool, xrate, for working with reversible, irreversible or parametric substitution models combined with stochastic context-free grammars. xrate efficiently estimates maximum-likelihood parameters and phylogenetic trees using a novel "phylo-EM" algorithm that we describe. The grammar is specified in an external configuration file, allowing users to design new grammars, estimate rate parameters from training data and annotate multiple sequence alignments without the need to recompile code from source. We have used xrate to measure codon substitution rates and predict protein and RNA secondary structures. CONCLUSION: Our results demonstrate that xrate estimates biologically meaningful rates and makes predictions whose accuracy is comparable to that of more specialized tools.

Algorithms↗

Influence of cyclical mechanical strain on extracellular matrix gene expression in human lamina cribrosa cells in vitro.

PURPOSE: The mechanical effect of raised intraocular pressure is a recognised stimulus for optic neuropathy in primary open angle glaucoma (POAG). Characteristic extracellular matrix (ECM) remodelling accompanies axonal damage in the lamina cribrosa (LC) of the optic nerve head in POAG. Glial cells in the lamina cribrosa may play a role in this process but the precise cellular responses to mechanical forces in this region are unknown. The authors examined global gene expression profiles in lamina cribrosa cells exposed to cyclical mechanical stretch, with an emphasis on ECM genes. METHODS: Glial fibrillary acid protein negative primary LC cells were generated from the optic nerve head tissue of three normal human donors. Confluent cell passages (n=4) were exposed to 15% stretch at 1 Hz or static conditions for 24 h using the Flexercell system. Gene expression was assessed using Affymetrix U133A microarrays with pooled RNA. Expression levels were normalized using robust multi-chip average (RMA). Expression data was annotated using NIH DAVID software. ECM-related gene expression was validated in an independent experiment using quantitative real-time PCR and protein synthesis was measured using ELISA and immunohistochemistry. RESULTS: Compared with static controls, 805 genes were upregulated and 644 were downregulated by +/-1.5 fold in stretched LC cells. Gene ontologies included ECM, cell proliferation, growth factor activity, and signal transduction. Differentially expressed ECM genes included elastin, collagens (IV, VI, VIII, IX), thrombospondin 1, perlecan, and lysl oxidase. Quantitative PCR demonstrated that the expression of TGF-beta2, BMP-7, elastin, collagen VI, biglycan, versican, EMMPRIN, VEGF, and thrombomodulin were reproducible and consistent with the microarray data. VEGF and TGF-beta2 protein levels were also significantly (p<0.05) increased in stretched cell media supernatants. Immunohistochemistry demonstrated increased EMMPRIN (an extracellular matrix metalloproteinase inducer) protein in human POAG optic nerve head tissue compared to nonglaucomatous controls. CONCLUSIONS: These findings demonstrate that LC cells respond to mechanical stimuli in vitro by transcription of several components and modulators of the ECM. Some of the upregulated ECM genes identified are novel in the context of glaucomatous optic neuropathy (biglycan, versican, EMMPRIN, and BMP-7). The LC cell may represent both an important pro-fibrotic cell type in the optic nerve head and an attractive target for novel therapeutic intervention in POAG.

Cells, Cultured↗

Synteny plot quality control with SyntenyQC.

SUMMARY: SyntenyQC is a data pre-processing tool for the construction of synteny plots. It supports genomic data collection, annotation and dereplication to facilitate (and in some cases fundamentally enable) the construction of informative synteny plots. AVAILABILITY AND IMPLEMENTATION: SyntenyQC is a command line app developed using Python version 3.10 and tested using pytest. SyntenyQC is available on PyPI (https://pypi.org/project/SyntenyQC) under the MIT License, along with a detailed user tutorial. Package tests can be viewed at https://github.com/Tim-Kirkwood/SyntenyQC.

Synteny↗