Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

ECG analysis for sleep apnea detection.

OBJECTIVES: The objective of our study was to find out whether obstructive sleep apnea (OSA) may be detected on ECGs recorded during sleep. METHODS: We have analyzed 70 eight-hour single-channel ECG recordings taken at polysomnographia. The 70 data sets were annotated for definition of regular sleep and phases with sleep apnea. From the 70 data sets, 35 have been used as a learning set. Our analysis is based on spectral components of heart rate variability. Frequency analysis was performed using Fourier and wavelet transformation with appropriate application of the Hilbert transform. Classification is based on four frequency bands: ULF band (0-0.013 Hz), VLF band (0.013-0.0375 Hz), LF band (0.0375-0.06 Hz) and the HF band (0.17-0.28 Hz). Linear discriminant functions were applied using mainly spectral components derived from the records. Classification of cases was based on three variables. RESULTS: For the Testing Set, a sensitivity (Se) for apnea of 92.3% at a specificity (Sp) of 94.6% was achieved. For the minutes allocation on the Learning Set Se was 90.8% at Sp 92.7%. CONCLUSION: ECG analysis is useful for the detection of sleep apnea and may help to differentiate causes of cardiac arrhythmias.

Electrocardiography↗

EST-PAC a web package for EST annotation and protein sequence prediction.

With the decreasing cost of DNA sequencing technology and the vast diversity of biological resources, researchers increasingly face the basic challenge of annotating a larger number of expressed sequences tags (EST) from a variety of species. This typically consists of a series of repetitive tasks, which should be automated and easy to use. The results of these annotation tasks need to be stored and organized in a consistent way. All these operations should be self-installing, platform independent, easy to customize and amenable to using distributed bioinformatics resources available on the Internet. In order to address these issues, we present EST-PAC a web oriented multi-platform software package for expressed sequences tag (EST) annotation. EST-PAC provides a solution for the administration of EST and protein sequence annotations accessible through a web interface. Three aspects of EST annotation are automated: 1) searching local or remote biological databases for sequence similarities using Blast services, 2) predicting protein coding sequence from EST data and, 3) annotating predicted protein sequences with functional domain predictions. In practice, EST-PAC integrates the BLASTALL suite, EST-Scan2 and HMMER in a relational database system accessible through a simple web interface. EST-PAC also takes advantage of the relational database to allow consistent storage, powerful queries of results and, management of the annotation process. The system allows users to customize annotation strategies and provides an open-source data-management environment for research and education in bioinformatics.

Journal Article↗

GenBank.

The GenBank sequence database incorporates DNA sequences from all available public sources, primarily through the direct submission of sequence data from authors and from large-scale sequencing projects. Data exchange with the EMBL Data Library and the DNA Data Bank of Japan helps ensure comprehensive coverage. GenBank continues to focus on quality control and annotation while expanding data coverage and retrieval services. An integrated retrieval system, known asEntrez, incorporates data from the major DNA and protein sequence databases, along with genome maps and protein structure information. MEDLINE abstracts from published articles describing the sequences are also included as an additional source of biological annotation. Sequence similarity searching is offered through the BLAST family of programs. All of NCBI's services are offered through the World Wide Web. In addition, there are specialized server/client versions as well as FTP and e-mail server access.

Amino Acid Sequence↗

Genome annotation: which tools do we have for it?

Genome data have to be converted into knowledge to be useful to biologists. Many valuable computational tools have already been developed to help annotation of plant genome sequences, and these may be improved further, for example by identification of more gene regulatory elements. The lack of a standard computer-assisted annotation platform for eukaryotic genomes remains major bottle-neck.

Arabidopsis↗

Local correlation of expression profiles with gene annotations--proof of concept for a general conciliatory method.

MOTIVATION: Integrated analysis of expression data and gene ontology annotations is a prime example of biological data that need co-explanatory interpretation. This particular application is used to validate a new method for integrated analysis of varied biological information. RESULTS: The proposed method consists of determining local correlation coefficients and the corresponding P-values calculated per biological entity. This measure considers the combined intensity and significance of the agreement or disagreement, between two data sources about the same biological entity. The method is applied to the integrated analysis of gene expression and annotation of two gene sets, one from yeast and other from mouse. The potential of the method to generate accurate mechanistic hypothesis is also demonstrated. Specially, negative correlation results pose a new kind of biological hypothesis. Method performance was compared with annotation enrichment methods, and optimal conditions for the superiority of local correlation results are discussed.

Algorithms↗

Gene expression informatics--it's all in your mine.

Technologies for whole-genome RNA expression studies are becoming increasingly reliable and accessible. However, universal standards to make the data more suitable for comparative analysis and for inter-operability with other information resources have yet to emerge. Improved access to large electronic data sets, reliable and consistent annotation and effective tools for 'data mining' are critical. Analysis methods that exploit large data warehouses of gene expression experiments will be necessary to realize the full potential of this technology.

Animals↗

Molecular classification and molecular genetics of human lung cancers.

Recent advances in the molecular classification of lung carcinomas and the identification of causative genetic alterations will likely lead to improvements in the diagnosis and treatment of patients with lung cancer. It is now possible to identify gene expression profiles that associate with patient outcome in lung carcinomas, in particular adenocarcinoma. Furthermore, patient survival has been shown to correlate with lung cancer oligonucleotide microarray expression profiles. Large-scale microarray technology may allow for the identification of useful biomarkers for early cancer detection. Oligonucleotide microarray data can be optimized by relating them to protein expression levels in tissue microarrays, by annotation with mutational data, and with results of testing for post-translational modification of cellular proteins. These data may be useful in tailoring chemotherapeutic protocols to individual tumors and identifying new targets for therapeutic intervention.

Adenocarcinoma↗

GenBank.

The GenBank nucleotide sequence database now contains sequence data and associated annotation corresponding to 56,000,000 nucleotides in 45,000 entries. The input stream of data coming into the database has largely been shifted to direct submissions from the scientific community on electronic media. The data have been installed in a relational database management system and are made available in this form through on-line access, and through various network and off-line computer-readable media. In addition, GenBank provides the U.S. distribution center for the BIOSCI electronic bulletin board service.

Base Sequence↗

Pre-Meta: priors-augmented retrieval for LLM-based metadata generation.

MOTIVATION: While high-throughput sequencing technologies have dramatically accelerated genomic data generation, the manual processes required for dataset annotation and metadata creation impede the efficient discovery and publication of these resources across disparate public repositories. Large language models (LLMs) have the potential to streamline dataset profiling and discovery. However, their current limitations in generalizing across specialized knowledge domains, particularly in fields such as biomedical genomics, prevent them from fully realizing this potential. This article presents Pre-Meta, an LLM-agnostic and domain-independent data annotation pipeline with an enriched retrieval procedure that leverages related priors-such as pre-generated metadata tags and ontologies-as auxiliary information to improve the accuracy of automated metadata generation. RESULTS: Validated using five selected metadata fields sampled across 1500 papers, the Pre-Meta assisted annotation experiment-without finetuning and prompt optimization-demonstrates a systemic improvement in the annotation task: shown through a 23%, 72%, and 75% accuracy gain from conventional RAG adoptions of GPT-4o mini, Llama 8B, and Mistral 7B respectively. AVAILABILITY AND IMPLEMENTATION: The code, data access, and scripts are available at: https://github.com/SINTEF-SE/LLMDap.

Metadata↗

BBP: Brucella genome annotation with literature mining and curation.

BACKGROUND: Brucella species are Gram-negative, facultative intracellular bacteria that cause brucellosis in humans and animals. Sequences of four Brucella genomes have been published, and various Brucella gene and genome data and analysis resources exist. A web gateway to integrate these resources will greatly facilitate Brucella research. Brucella genome data in current databases is largely derived from computational analysis without experimental validation typically found in peer-reviewed publications. It is partially due to the lack of a literature mining and curation system able to efficiently incorporate the large amount of literature data into genome annotation. It is further hypothesized that literature-based Brucella gene annotation would increase understanding of complicated Brucella pathogenesis mechanisms. RESULTS: The Brucella Bioinformatics Portal (BBP) is developed to integrate existing Brucella genome data and analysis tools with literature mining and curation. The BBP InterBru database and Brucella Genome Browser allow users to search and analyze genes of 4 currently available Brucella genomes and link to more than 20 existing databases and analysis programs. Brucella literature publications in PubMed are extracted and can be searched by a TextPresso-powered natural language processing method, a MeSH browser, a keywords search, and an automatic literature update service. To efficiently annotate Brucella genes using the large amount of literature publications, a literature mining and curation system coined Limix is developed to integrate computational literature mining methods with a PubSearch-powered manual curation and management system. The Limix system is used to quickly find and confirm 107 Brucella gene mutations including 75 genes shown to be essential for Brucella virulence. The 75 genes are further clustered using COG. In addition, 62 Brucella genetic interactions are extracted from literature publications. These results make possible more comprehensive investigation of Brucella pathogenesis. Other BBP features include publication email alert service, Brucella researchers' contact database, and discussion forum. CONCLUSION: BBP is a gateway for Brucella researchers to search, analyze, and curate Brucella genome data originated from public databases and literature. Brucella gene mutations and genetic interactions are annotated using Limix leading to better understanding of Brucella pathogenesis.

Algorithms↗

Application of InChI to curate, index, and query 3-D structures.

The HIV structural database (HIVSDB) is a comprehensive collection of the structures of HIV protease, both of unliganded enzyme and of its inhibitor complexes. It contains abstracts and crystallographic data such as inhibitor and protein coordinates for 248 data sets, of which only 141 are from the Protein Data Bank (PDB). Efficient annotation, indexing, and querying of the inhibitor data is crucial for their effective use for technological and industrial applications. The application of IUPAC International Chemical Identifier (InChI) to index, curate, and query inhibitor structures HIVSDB is described.

Abstracting and Indexing↗

DRAGON: Database Referencing of Array Genes Online.

UNLABELLED: "Database Referencing of Array Genes ONline" or "DRAGON" is a web-accessible database that aids in the analysis of differential gene expression data as a biological annotation tool. Users of DRAGON can submit data sets containing large lists of genes and then choose particular characteristics that DRAGON supplies for all genes on the list rapidly and simultaneously. AVAILABILITY: The DRAGON database is available for queries on the DRAGON web site www.kennedykrieger.org/pevsnerlab/dragon.htm. CONTACT: pevsner@kennedykrieger.org or cbouton@jhmi.edu

Computational Biology↗

HmtDB, a human mitochondrial genomic resource based on variability studies supporting population genetics and biomedical research.

BACKGROUND: Population genetics studies based on the analysis of mtDNA and mitochondrial disease studies have produced a huge quantity of sequence data and related information. These data are at present worldwide distributed in differently organised databases and web sites not well integrated among them. Moreover it is not generally possible for the user to submit and contemporarily analyse its own data comparing them with the content of a given database, both for population genetics and mitochondrial disease data. RESULTS: HmtDB is a well-integrated web-based human mitochondrial bioinformatic resource aimed at supporting population genetics and mitochondrial disease studies, thanks to a new approach based on site-specific nucleotide and aminoacid variability estimation. HmtDB consists of a database of Human Mitochondrial Genomes, annotated with population data, and a set of bioinformatic tools, able to produce site-specific variability data and to automatically characterize newly sequenced human mitochondrial genomes. A query system for the retrieval of genomes and a web submission tool for the annotation of new genomes have been designed and will soon be implemented. The first release contains 1255 fully annotated human mitochondrial genomes. Nucleotide site-specific variability data and multialigned genomes can be downloaded. Intra-human and inter-species aminoacid variability data estimated on the 13 coding for proteins genes of the 1255 human genomes and 60 mammalian species are also available. HmtDB is freely available, upon registration, at http://www.hmdb.uniba.it. CONCLUSION: The HmtDB project will contribute towards completing and/or refining haplogroup classification and revealing the real pathogenic potential of mitochondrial mutations, on the basis of variability estimation.

Computational Biology↗

ABC: software for interactive browsing of genomic multiple sequence alignment data.

BACKGROUND: Alignment and comparison of related genome sequences is a powerful method to identify regions likely to contain functional elements. Such analyses are data intensive, requiring the inclusion of genomic multiple sequence alignments, sequence annotations, and scores describing regional attributes of columns in the alignment. Visualization and browsing of results can be difficult, and there are currently limited software options for performing this task. RESULTS: The Application for Browsing Constraints (ABC) is interactive Java software for intuitive and efficient exploration of multiple sequence alignments and data typically associated with alignments. It is used to move quickly from a summary view of the entire alignment via arbitrary levels of resolution to individual alignment columns. It allows for the simultaneous display of quantitative data, (e.g., sequence similarity or evolutionary rates) and annotation data (e.g. the locations of genes, repeats, and constrained elements). It can be used to facilitate basic comparative sequence tasks, such as export of data in plain-text formats, visualization of phylogenetic trees, and generation of alignment summary graphics. CONCLUSIONS: The ABC is a lightweight, stand-alone, and flexible graphical user interface for browsing genomic multiple sequence alignments of specific loci, up to hundreds of kilobases or a few megabases in length. It is coded in Java for cross-platform use and the program and source code are freely available under the General Public License. Documentation and a sample data set are also available http://mendel.stanford.edu/sidowlab/downloads.html.

Animals↗

Comparative promoter analysis in vertebrate genomes with the CORG workbench.

CORG is a versatile web-based workbench for comparative promoter analysis in vertebrate model organisms. Two kinds of information are explicitly considered in the automated annotation process. First, local conservation patterns in upstream regions of homologous genes: These phylogenetic footprints are likely to stem from sequence elements that are under selective pressure. The CORG pipeline detects and exploits patterns of local similarity to annotate promoter regions. Second, experimental data on transcription start sites: exon positions and DNA binding site descriptions complete the promoter annotation. These data are made available via an interactive web portal. Individual promoter studies are supported by a JAVA applet that supplies all data down to the nucleotide level.

Animals↗

Bioinformatics strategies for translating genome-wide expression analyses into clinically useful cancer markers.

The DNA microarray has revolutionized cancer research. Now, scientists can obtain a genome-wide perspective of cancer gene expression. One potential application of this technology is the discovery of novel cancer biomarkers for more accurate diagnosis and prognosis, and potentially for the earlier detection of disease or the monitoring of treatment effectiveness. Because microarray experiments generate a tremendous amount of data and because the number of laboratories generating microarray data is rapidly growing, new bioinformatics strategies that promote the maximum utilization of such data are necessary. Here, we describe a method to validate multiple microarray data sets, a Web-based cancer microarray database for biomarker discovery, and methods for integrating gene ontology annotations with microarray data to improve candidate biomarker selection.

Biomarkers, Tumor↗

Proteomic approaches to studying drug targets and resistance in Plasmodium.

Ever increasing drug resistance by Plasmodium falciparum, the most virulent of human malaria parasites, is creating new challenges in malaria chemotherapy. The entire genome sequences of P. falciparum and the rodent malaria parasite, P. yoelii yoelii are now available. Extensive genome sequence data from other Plasmodium species including another important human malaria parasite, P. vivax are also available. Powerful research techniques coupled to genomic resources are needed to help identify new drug and vaccine targets against malaria. Applied to Plasmodium, proteomics combines high-resolution protein or peptide separation with mass spectrometry and computer software to rapidly identify large numbers of proteins expressed from various stages of parasite development. Proteomic methods can be applied to study sub-cellular localization, cell function, organelle composition, changes in protein expression patterns in response to drug exposure, drug-protein binding and validation of data from genomic annotation and transcript expression studies. Recent high-throughput proteomic approaches have provided a wealth of protein expression data on P. falciparum, while smaller-scale studies examining specific drug-related hypotheses are also appearing. Of particular interest is the study of mechanisms of action and resistance of drugs such as the quinolines, whose targets currently may not be predictable from genomic data. Coupling the Plasmodium sequence data with bioinformatics, proteomics and RNA transcript expression profiling opens unprecedented opportunities for exploring new malaria control strategies. This review will focus on pharmacological research in malaria and other intracellular parasites using proteomic techniques, emphasizing resources and strategies available for Plasmodium.

Animals↗

Open-source toolkit for simple XML annotation.

Use of Extensible Markup Language (XML) is increasingly prevalent among medical informatics projects. Many of these projects involve, at some point, the interaction between a researcher and specialized XML documents for the purpose of annotating the XML data. We offer a simple toolkit to assist these researchers. Our solution is a simple, yet fully functional, annotation system that can easily be adapted to the needs of the researcher. All of the materials for this toolkit are freely available.

Algorithms↗