Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Investigating semantic similarity measures across the Gene Ontology: the relationship between sequence and annotation.

MOTIVATION: Many bioinformatics data resources not only hold data in the form of sequences, but also as annotation. In the majority of cases, annotation is written as scientific natural language: this is suitable for humans, but not particularly useful for machine processing. Ontologies offer a mechanism by which knowledge can be represented in a form capable of such processing. In this paper we investigate the use of ontological annotation to measure the similarities in knowledge content or 'semantic similarity' between entries in a data resource. These allow a bioinformatician to perform a similarity measure over annotation in an analogous manner to those performed over sequences. A measure of semantic similarity for the knowledge component of bioinformatics resources should afford a biologist a new tool in their repertoire of analyses. RESULTS: We present the results from experiments that investigate the validity of using semantic similarity by comparison with sequence similarity. We show a simple extension that enables a semantic search of the knowledge held within sequence databases. AVAILABILITY: Software available from http://www.russet.org.uk.

Artificial Intelligence↗

ArrayExpress--a public repository for microarray gene expression data at the EBI.

ArrayExpress is a new public database of microarray gene expression data at the EBI, which is a generic gene expression database designed to hold data from all microarray platforms. ArrayExpress uses the annotation standard Minimum Information About a Microarray Experiment (MIAME) and the associated XML data exchange format Microarray Gene Expression Markup Language (MAGE-ML) and it is designed to store well annotated data in a structured way. The ArrayExpress infrastructure consists of the database itself, data submissions in MAGE-ML format or via an online submission tool MIAMExpress, online database query interface, and the Expression Profiler online analysis tool. ArrayExpress accepts three types of submission, arrays, experiments and protocols, each of these is assigned an accession number. Help on data submission and annotation is provided by the curation team. The database can be queried on parameters such as author, laboratory, organism, experiment or array types. With an increasing number of organisations adopting MAGE-ML standard, the volume of submissions to ArrayExpress is increasing rapidly. The database can be accessed at http://www.ebi.ac.uk/arrayexpress.

Animals↗

Two-dimensional gel protein database of Saccharomyces cerevisiae.

With the systematic sequencing of the yeast genome, yeast biology has entered a new era where novel challenges have to be faced. One challenge is the identification of the function of the several hundred novel genes discovered by genome sequencing. Another is to understand how all yeast genes act in concert to ensure and maintain cell organization. Two-dimensional (2-D) gel electrophoresis is the technique of choice to take up these challenges because it provides the opportunity of obtaining an overall view of genome expression. In prospect of these studies we have undertaken the construction of a yeast 2-D gel protein database that contains information on polypeptides of the yeast protein map. In this paper we report the information presently contained in this database. The reported information includes the identification of 250 protein spots and the characterization of polypeptides corresponding to N-terminal acetylated proteins, mitochondrial proteins, glucose-repressed proteins, heat shock induced proteins and proteins encoded by intron-containing genes. In all, 600 spots are annotated. These data can be accessed on the Yeast Protein Map server through the World Wide Web network.

Computer Communication Networks↗

A PC-based generator of surface ECG potentials for computer electrocardiograph testing.

The system is composed of an electronic circuit, connected to a PC, whose outputs, starting from ECGs digitally collected by commercial interpretative electrocardiographs, simulate virtual patients' limb and chest electrode potentials. Appropriate software manages the D/A conversion and lines up the original short-term signal in a ring buffer to generate continuous ECG traces. The device also permits the addition of artifacts and/or baseline wanders/shifts on each lead separately. The system has been accurately tested and statistical indexes have been computed to quantify the reproduction accuracy analyzing, in the generated signal, both the errors induced on the fiducial point measurements and the capability to retain the diagnostic significance. The device integrated with an annotated ECG data base constitutes a reliable and powerful system to be used in the quality assurance testing of computer electrocardiographs.

Electrocardiography↗

Regulation of gene expression in flux balance models of metabolism.

Genome-scale metabolic networks can now be reconstructed based on annotated genomic data augmented with biochemical and physiological information about the organism. Mathematical analysis can be performed to assess the capabilities of these reconstructed networks. The constraints-based framework, with flux balance analysis (FBA), has been used successfully to predict time course of growth and by-product secretion, effects of mutation and knock-outs, and gene expression profiles. However, FBA leads to incorrect predictions in situations where regulatory effects are a dominant influence on the behavior of the organism. Thus, there is a need to include regulatory events within FBA to broaden its scope and predictive capabilities. Here we represent transcriptional regulatory events as time-dependent constraints on the capabilities of a reconstructed metabolic network to further constrain the space of possible network functions. Using a simplified metabolic/regulatory network, growth is simulated under various conditions to illustrate systemic effects such as catabolite repression, the aerobic/anaerobic diauxic shift and amino acid biosynthesis pathway repression. The incorporation of transcriptional regulatory events in FBA enables us to interpret, analyse and predict the effects of transcriptional regulation on cellular metabolism at the systemic level.

Amino Acids↗

A computer model of intracranial pressure dynamics during traumatic brain injury that explicitly models fluid flows and volumes.

A model of intracranial pressure (ICP) dynamics that uses fluid volumes as primary state variables is presented, along with clinical data for two subjects with elevated ICP. The data includes annotations to indicate the precise timing of clinical changes in cerebral spinal fluid drainage, head of bed elevation, and minute ventilation. The response to changes in the clinical parameters was used to calibrate the model to correspond to specific subjects by estimating values for key characteristics such as hematoma volume and CSF uptake resistance. The error in mean ICP predicted by the model was less than 2 mmHg when cerebral spinal fluid is drained and the head of bed elevation was increased. The error in mean ICP predicted by the model exceeded 5 mmHg during an episode when the head of bed was decreased and also during a reduction in minute ventilation. The estimated values for hematoma volume and other subject characteristics were plausible but could not be verified empirically.

Blood Pressure↗

The Macrostomum lignano EST database as a molecular resource for studying platyhelminth development and phylogeny.

We report the development of an Expressed Sequence Tag (EST) resource for the flatworm Macrostomum lignano. This taxon is of interest due to its basal placement within the flatworms. As such, it provides a useful comparative model for understanding the development of neural and sensory organization. It was anticipated on the basis of previous studies [e.g., Sánchez-Alvarado et al., Development, 129:5659-5665, (2002)] that a wide range of developmental markers would be expressed in later-stage macrostomids, and this proved to be the case, permitting recovery of a range of gene sequences important in development. To this end, an adult Macrostomum cDNA library was generated and 7,680 Macrostomum ESTs were sequenced from the 5' end. In addition, 1,536 of these aforementioned sequences were sequenced from the 3' end. Of the roughly 5,416 non-redundant sequences identified, 68% are similar to previously reported genes of known function. In addition, nearly 100 specific clones were obtained with potential neural and sensory function. From these data, an annotated searchable database of the Macrostomum EST collection has been made available on the web. A major objective was to obtain genes that would allow reconstruction of embryogenesis, and in particular neurogenesis, in a basal platyhelminth. The sequences recovered will serve as probes with which the origin and morphogenesis of lineages and tissues can be followed. To this end, we demonstrate a protocol for combined immunohistochemistry and in situ hybridization labeling in juvenile Macrostomum, employing homologs of lin11/lim1 and six3/optix. Expression of these genes is shown in the context of the neuropile/muscle system.

Animals↗

Automated segmentation of the left ventricle in cardiac MRI.

We present a fully automated deformable model technique for myocardium segmentation in 3D MRI. Loss of signal due to blood flow, partial volume effects and significant variation of surface grey value appearance make this a difficult problem. We integrate various sources of prior knowledge learned from annotated image data into a deformable model. Inter-individual shape variation is represented by a statistical point distribution model, and the spatial relationship of the epi- and endocardium is modeled by adapting two coupled triangular surface meshes. To robustly accommodate variation of grey value appearance around the myocardiac surface, a prior parametric spatially varying feature model is established by classification of grey value surface profiles. Quantitative validation of 121 3D MRI datasets in end-diastolic (end-systolic) phase demonstrates accuracy and robustness, with 2.45 mm (2.84 mm) mean deviation from manual segmentation.

Algorithms↗

The Elegance of the MicroRNAs: A Neuronal Perspective.

As knowledge of microRNAs (miRNA) grows from a compendium of sequences to annotated functional data it has become increasingly clear that a highly significant segment of regulatory biology depends on these approximately 22 nucleotide noncoding transcripts. The expression of many miRNAs in the nervous system, some with a high degree of temporal and spatial specificity, suggests that understanding miRNAs in the nervous system will yield rewarding neurobiological insights. High on the list of insights that microRNAs promise is a deeper understanding of the remarkable cellular diversity found among neurons. This review examines the interface between an emerging biology of miRNAs and their role in nervous systems.

Animals↗

Moonlighting vacuolar protease: multiple jobs for a busy protein.

In this Genomics Era with a wealth of annotated sequence data, it is easy to pigeonhole a protein into a particular function. However, Noa Matarasso et al. recently found a vacuolar protease that can also function as a transcription factor. This work illustrates that a protein can serve multiple roles in a cell, raising intriguing questions as to the extent that genomic information can be deciphered de novo.

Ethylenes↗

High-resolution RH map of horse chromosome 22 reveals a putative ancestral vertebrate chromosome.

High-resolution gene maps of individual equine chromosomes are essential to identify genes governing traits of economic importance in the horse. In pursuit of this goal we herein report the generation of a dense map of horse chromosome 22 (ECA22) comprising 83 markers, of which 52 represent specific genes and 31 are microsatellites. The map spans 831 cR over an estimated 64 Mb of physical length of the chromosome, thus providing markers at approximately 770 kb or 10 cR intervals. Overall, the resolution of the map is to date the densest in the horse and is the highest for any of the domesticated animal species for which annotated sequence data are not yet available. Comparative analysis showed that ECA22 shares remarkable conservation of gene order along the entire length of dog chromosome 24, something not yet found for an autosome in evolutionarily diverged species. Comparison with human, mouse, and rat homologues shows that ECA22 can be traced as two conserved linkage blocks, each related to individual arms of the human homologue-HSA20. Extending the comparison to the chicken genome showed that one of the ECA22 blocks that corresponds to HSA20q shares synteny conservation with chicken chromosome 20, suggesting the segment to be ancestral in mammals and birds.

Animals↗

Parasite genome initiatives.

During 1993-1994, scientists from developing and developed countries planned and initiated a number of parasite genome projects and several consortiums for the mapping and sequencing of these medium-sized genomes were established, often based on already ongoing scientific collaborations. Financial and other support came from WHO/TDR, Wellcome Trust and other funding agencies. Thus, the genomes of Plasmodium falciparum, Schistosoma mansoni, Trypanosoma cruzi, Leishmania major, Trypanosoma brucei, Brugia malayi and other pathogenic nematodes are now under study. From an initial phase of network formation, mapping efforts and resource building (EST, GSS, phage, cosmid, BAC and YAC library constructions), sequencing was initiated in gene discovery projects but soon also on a small chromosome, and now on a fully fledged genome scale. Proteomics, functional analysis, genetic manipulation and microarray analysis are ongoing to different degrees in the respective genome initiatives, and as the funding for the whole genome sequencing becomes secured, most of the participating laboratories, apart from larger sequencing centres, become oriented to post-genomics. Bioinformatics networks are being expanded, including in developing countries, for data mining, annotation and in-depth analysis.

Animals↗

An integrated Arabidopsis annotation database for Affymetrix Genechip data analysis, and tools for regulatory motif searches.

Genome-scale sequencing projects have provided the essential information required for the construction of entire genome chips or microarrays for RNA expression studies. The Arabidopsis and rice genomes have been sequenced and whole-genome oligonucleotide arrays are being manufactured. These should soon become available to researchers. Expression studies using genomic-scale expression arrays are providing us with a vast quantity of information at a rapid pace. The rate-limiting step in this type of experiments is not the data generation step but rather the data analysis component of experiments. We report improvements that should facilitate the analysis of Affymetrix Genechip expression data.

Arabidopsis↗

Rutabaga by any other name: extracting biological names.

As the pace of biological research accelerates, biologists are becoming increasingly reliant on computers to manage the information explosion. Biologists communicate their research findings by relying on precise biological terms; these terms then provide indices into the literature and across the growing number of biological databases. This article examines emerging techniques to access biological resources through extraction of entity names and relations among them. Information extraction has been an active area of research in natural language processing and there are promising results for information extraction applied to news stories, e.g., balanced precision and recall in the 93-95% range for identifying person, organization and location names. But these results do not seem to transfer directly to biological names, where results remain in the 75-80% range. Multiple factors may be involved, including absence of shared training and test sets for rigorous measures of progress, lack of annotated training data specific to biological tasks, pervasive ambiguity of terms, frequent introduction of new terms, and a mismatch between evaluation tasks as defined for news and real biological problems. We present evidence from a simple lexical matching exercise that illustrates some specific problems encountered when identifying biological names. We conclude by outlining a research agenda to raise performance of named entity tagging to a level where it can be used to perform tasks of biological importance.

Abstracting and Indexing↗

Modeling hybridoma cell metabolism using a generic genome-scale metabolic model of Mus musculus.

The reconstructed cellular metabolic network of Mus musculus, based on annotated genomic data, pathway databases, and currently available biochemical and physiological information, is presented. Although incomplete, it represents the first attempt to collect and characterize the metabolic network of a mammalian cell on the basis of genomic data. The reaction network is generic in nature and attempts to capture the carbon, energy, and nitrogen metabolism of the cell. The metabolic reactions were compartmentalized between the cytosol and the mitochondria, including transport reactions between the compartments and the extracellular medium. The reaction list consists of 872 internal metabolites involved in a total of 1220 reactions, whereof 473 relate to known open reading frames. Initial in silico analysis of the reconstructed model is presented.

Amino Acids↗

PubMind: literature-based genetic variant extraction and functional annotation using large language models.

Biomedical literature contains extensive functional knowledge on genetic variants, but much remains inaccessible in unstructured text. Existing resources such as ClinVar and HGMD remain limited by coverage, submission bias, update frequency, and sparse annotation. We develop PubMind, an artificial intelligence (AI) framework that uses large language models (LLMs) to triage and extract variant-function-disease associations and supporting evidence from biomedical text. PubMind captures single-nucleotide, copy-number, structural, and gene-fusion variants, and normalizes records to genomic and transcriptomic coordinates. Benchmarking shows >90% accuracy for variant recognition and 99% precision for disease extraction. Applied to >41 million PubMed abstracts and >5 million full-text articles, PubMind generates PubMind-DB, a database of ~1.3 million unique variants with contextual annotations, accessible via web interface and API. Only ~10% of PubMind variants overlap with ClinVar, and >80% of them show concordant pathogenicity labels. PubMind transforms unstructured biomedical text into structured genomic knowledge, advancing variant interpretation for precision medicine.

Large Language Models↗

A systematic review of clinical research addressing the prevalence, aetiology, diagnosis, prognosis and therapy of otitis media in Australian Aboriginal children.

The objective of this review was to systemically identify and summarize all the clinically relevant evidence available from studies addressing the prevalence, aetiology, diagnosis, prognosis and therapy of otitis media in Australian Aboriginal children. Electronic searching of Medline, the Australian Medical Index and the Aboriginal and Torres Strait Islander Health Bibliographic Index was performed. This was supplemented by hand searching the Menzies School of Health Research otitis media collection, the Aboriginal and Torres Strait Islander Health Information Bulletin and Aboriginal Health: an annotated bibliography. Data were extracted and placed in a series of evidence tables relevant to clinical practice. There were 59 studies that met the inclusion criteria. The majority were surveys, and only 19 addressed diagnosis, prognosis or therapy. Severe otitis media in rural Aboriginal children does not occur in isolation but as part of a spectrum of chronic bacterial infections of the respiratory tract. Although the aspects of poverty that result in this condition remain to be clarified, exposure to other young children with chronic nasal discharge is likely to be important. Whilst there is a considerable amount of literature on otitis media in Australian Aboriginal children, the number of studies most relevant to improving health outcomes is small. A systematic approach to disease surveillance, diagnosis, and application of medical interventions is required urgently. Future medical research should be concerned with the evaluation of interventions and the generalisabilty of studies from different populations.

Adolescent↗

How not to be seen: predicting unseen enzyme functions using contrastive learning.

MOTIVATION: Predicting enzyme function from its sequence is still an unsolved problem in the life sciences. Moreover, with the explosion of annotated genome data, we are inundated with potential enzymatic sequences that have not yet been biochemically characterized. While it is not possible to assign a not-yet-existing label to such a sequence, there is high value in placing the sequence as accurately as possible in known function space. Doing so can help provide more accurate falsifiable hypotheses for experimentalists wishing to characterize enzymes from specific functional families. RESULTS: Here we present a contrastive learning algorithm for predicting enzyme function from sequence. Our method, EnzPlacer, predicts the third, second, and first EC numbers for a protein whose fourth EC number is not in the training corpus. This novel prediction mechanism accurately places a protein sequence within a narrowed-down functional context, even if the precise function remains unknown. AVAILABILITY AND IMPLEMENTATION: EnzPlacer and data is available at https://github.com/drxiangma/EnzPlacer under a GPL3 license.

Enzymes↗