Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

Novel gene and gene model detection using a whole genome open reading frame analysis in proteomics.

BACKGROUND: Defining the location of genes and the precise nature of gene products remains a fundamental challenge in genome annotation. Interrogating tandem mass spectrometry data using genomic sequence provides an unbiased method to identify novel translation products. A six-frame translation of the entire human genome was used as the query database to search for novel blood proteins in the data from the Human Proteome Organization Plasma Proteome Project. Because this target database is orders of magnitude larger than the databases traditionally employed in tandem mass spectra analysis, careful attention to significance testing is required. Confidence of identification is assessed using our previously described Poisson statistic, which estimates the significance of multi-peptide identifications incorporating the length of the matching sequence, number of spectra searched and size of the target sequence database. RESULTS: Applying a false discovery rate threshold of 0.05, we identified 282 significant open reading frames, each containing two or more peptide matches. There were 627 novel peptides associated with these open reading frames that mapped to a unique genomic coordinate placed within the start/stop points of previously annotated genes. These peptides matched 1,110 distinct tandem MS spectra. Peptides fell into four categories based upon where their genomic coordinates placed them relative to annotated exons within the parent gene. CONCLUSION: This work provides evidence for novel alternative splice variants in many previously annotated genes. These findings suggest that annotation of the genome is not yet complete and that proteomics has the potential to further add to our understanding of gene structures.

Alternative Splicing↗

Transcriptome analysis of the diseased intervertebral disc tissue in patients with spinal tuberculosis.

OBJECTIVE: To investigate the differential expression genes (DEGs) in spinal tuberculosis using transcriptomics, with the aim of identifying novel therapeutic targets and prognostic indicators for the clinical management of spinal tuberculosis. METHODS: Patients who visited the Department of Orthopedics at the Second Hospital, Lanzhou University from January 2021 to May 2023 were enrolled. Based on the inclusion and exclusion criteria, there were 5 patients in the test group and 5 patients in the control group. Total RNA was extracted and paired-end sequencing was conducted on the sequencing platform. After processing the sequencing data with clean reads and annotating the reference genome, FPKM normalization and differential expression analysis were performed. The DEGs and long non-coding RNAs (LncRNAs) were analyzed for Kyoto Encyclopedia of Genes and Genomes (KEGG) and Gene Ontology (GO) enrichment. The cis-regulation of differentially expressed mRNAs (DE mRNAs) by LncRNAs was predicted and analyzed to establish a co-expression network. RESULTS: This study identified 2366 DEGs, with 974 genes significantly upregulated and 1392 genes significantly downregulated. The upregulated genes are associated with cytokine-cytokine receptor interactions, tuberculosis, and TNF-α signaling pathways, primarily enriched in biological processes such as immunity and inflammation. The downregulated genes are related to muscle development, contraction, fungal defense response, and collagen metabolism processes. Analysis of LncRNAs from bone tuberculosis RNA-seq data detected a total of 3652 LncRNAs, with 356 significantly upregulated and 184 significantly downregulated. Further analysis identified 311 significantly different LncRNAs that could cis-regulate 777 target genes, enriched in pathways such as muscle contraction, inflammatory response, and immune response, closely related to bone tuberculosis. There are 51 genes enriched in the immune response pathway regulated by cis-acting LncRNAs. LncRNAs that regulate immune response-related genes, such as upregulated RP11-451G4.2, RP11-701P16.5, AC079767.4, AC017002.1, LINC01094, CTA-384D8.35, and AC092484.1, as well as downregulated RP11-2C24.7, may serve as potential prognostic and therapeutic targets. CONCLUSION: The DE mRNAs and LncRNAs in spinal tuberculosis are both associated with immune regulatory pathways. These pathways promote or inhibit the tuberculosis infection and development at the mechanistic level and play an important role in the process of tuberculosis transferring to bone tissue.

Humans↗

AceView: a comprehensive cDNA-supported gene and transcripts annotation.

BACKGROUND: Regions covering one percent of the genome, selected by ENCODE for extensive analysis, were annotated by the HAVANA/Gencode group with high quality transcripts, thus defining a benchmark. The ENCODE Genome Annotation Assessment Project (EGASP) competition aimed at reproducing Gencode and finding new genes. The organizers evaluated the protein predictions in depth. We present a complementary analysis of the mRNAs, including alternative transcript variants. RESULTS: We evaluate 25 gene tracks from the University of California Santa Cruz (UCSC) genome browser. We either distinguish or collapse the alternative splice variants, and compare the genomic coordinates of exons, introns and nucleotides. Whole mRNA models, seen as chains of introns, are sorted to find the best matching pairs, and compared so that each mRNA is used only once. At the mRNA level, AceView is by far the closest to Gencode: the vast majority of transcripts of the two methods, including alternative variants, are identical. At the protein level, however, due to a lack of experimental data, our predictions differ: Gencode annotates proteins in only 41% of the mRNAs whereas AceView does so in virtually all. We describe the driving principles of AceView, and how, by performing hand-supervised automatic annotation, we solve the combinatorial splicing problem and summarize all of GenBank, dbEST and RefSeq into a genome-wide non-redundant but comprehensive cDNA-supported transcriptome. AceView accuracy is now validated by Gencode. CONCLUSION: Relative to a consensus mRNA catalog constructed from all evidence-based annotations, Gencode and AceView have 81% and 84% sensitivity, and 74% and 73% specificity, respectively. This close agreement validates a richer view of the human transcriptome, with three to five times more transcripts than in UCSC Known Genes (sensitivity 28%), RefSeq (sensitivity 21%) or Ensembl (sensitivity 19%).

Computational Biology↗

Statistical Viewer: a tool to upload and integrate linkage and association data as plots displayed within the Ensembl genome browser.

BACKGROUND: To facilitate efficient selection and the prioritization of candidate complex disease susceptibility genes for association analysis, increasingly comprehensive annotation tools are essential to integrate, visualize and analyze vast quantities of disparate data generated by genomic screens, public human genome sequence annotation and ancillary biological databases. We have developed a plug-in package for Ensembl called "Statistical Viewer" that facilitates the analysis of genomic features and annotation in the regions of interest defined by linkage analysis. RESULTS: Statistical Viewer is an add-on package to the open-source Ensembl Genome Browser and Annotation System that displays disease study-specific linkage and/or association data as 2 dimensional plots in new panels in the context of Ensembl's Contig View and Cyto View pages. An enhanced upload server facilitates the upload of statistical data, as well as additional feature annotation to be displayed in DAS tracts, in the form of Excel Files. The Statistical View panel, drawn directly under the ideogram, illustrates lod score values for markers from a study of interest that are plotted against their position in base pairs. A module called "Get Map" easily converts the genetic locations of markers to genomic coordinates. The graph is placed under the corresponding ideogram features a synchronized vertical sliding selection box that is seamlessly integrated into Ensembl's Contig- and Cyto- View pages to choose the region to be displayed in Ensembl's "Overview" and "Detailed View" panels. To resolve Association and Fine mapping data plots, a "Detailed Statistic View" plot corresponding to the "Detailed View" may be displayed underneath. CONCLUSION: Features mapping to regions of linkage are accentuated when Statistic View is used in conjunction with the Distributed Annotation System (DAS) to display supplemental laboratory information such as differentially expressed disease genes in private data tracks. Statistic View is a novel and powerful visual feature that enhances Ensembl's utility as valuable resource for integrative genomic-based approaches to the identification of candidate disease susceptibility genes. At present there are no other tools that provide for the visualization of 2-dimensional plots of quantitative data scores against genomic coordinates in the context of a primary public genome annotation browser.

Chromosome Mapping↗

Evaluation of gene prediction software using a genomic data set: application to Arabidopsis thaliana sequences.

MOTIVATION: The annotation of the Arabidopsis thaliana genome remains a problem in terms of time and quality. To improve the annotation process, we want to choose the most appropriate tools to use inside a computer-assisted annotation platform. We therefore need evaluation of prediction programs with Arabidopsis sequences containing multiple genes. RESULTS: We have developed AraSet, a data set of contigs of validated genes, enabling the evaluation of multi-gene models for the Arabidopsis genome. Besides conventional metrics to evaluate gene prediction at the site and the exon levels, new measures were introduced for the prediction at the protein sequence level as well as for the evaluation of gene models. This evaluation method is of general interest and could apply to any new gene prediction software and to any eukaryotic genome. The GeneMark.hmm program appears to be the most accurate software at all three levels for the Arabidopsis genomic sequences. Gene modeling could be further improved by combination of prediction software. AVAILABILITY: The AraSet sequence set, the Perl programs and complementary results and notes are available at http://sphinx.rug.ac.be:8080/biocomp/napav/. CONTACT: Pierre.Rouze@gengenp.rug.ac.be.

Alternative Splicing↗

AgBase: a unified resource for functional analysis in agriculture.

Analysis of functional genomics (transcriptomics and proteomics) datasets is hindered in agricultural species because agricultural genome sequences have relatively poor structural and functional annotation. To facilitate systems biology in these species we have established the curated, web-accessible, public resource 'AgBase' (www.agbase.msstate.edu). We have improved the structural annotation of agriculturally important genomes by experimentally confirming the in vivo expression of electronically predicted proteins and by proteogenomic mapping. Proteogenomic data are available from the AgBase proteogenomics link. We contribute Gene Ontology (GO) annotations and we provide a two tier system of GO annotations for users. The 'GO Consortium' gene association file contains the most rigorous GO annotations based solely on experimental data. The 'Community' gene association file contains GO annotations based on expert community knowledge (annotations based directly from author statements and submitted annotations from the community) and annotations for predicted proteins. We have developed two tools for proteomics analysis and these are freely available on request. A suite of tools for analyzing functional genomics datasets using the GO is available online at the AgBase site. We encourage and publicly acknowledge GO annotations from researchers and provide an online mechanism for agricultural researchers to submit requests for GO annotations.

Agriculture↗

Insights into a dinoflagellate genome through expressed sequence tag analysis.

BACKGROUND: Dinoflagellates are important marine primary producers and grazers and cause toxic "red tides". These taxa are characterized by many unique features such as immense genomes, the absence of nucleosomes, and photosynthetic organelles (plastids) that have been gained and lost multiple times. We generated EST sequences from non-normalized and normalized cDNA libraries from a culture of the toxic species Alexandrium tamarense to elucidate dinoflagellate evolution. Previous analyses of these data have clarified plastid origin and here we study the gene content, annotate the ESTs, and analyze the genes that are putatively involved in DNA packaging. RESULTS: Approximately 20% of the 6,723 unique (11,171 total 3'-reads) ESTs data could be annotated using Blast searches against GenBank. Several putative dinoflagellate-specific mRNAs were identified, including one novel plastid protein. Dinoflagellate genes, similar to other eukaryotes, have a high GC-content that is reflected in the amino acid codon usage. Highly represented transcripts include histone-like (HLP) and luciferin binding proteins and several genes occur in families that encode nearly identical proteins. We also identified rare transcripts encoding a predicted protein highly similar to histone H2A.X. We speculate this histone may be retained for its role in DNA double-strand break repair. CONCLUSION: This is the most extensive collection to date of ESTs from a toxic dinoflagellate. These data will be instrumental to future research to understand the unique and complex cell biology of these organisms and for potentially identifying the genes involved in toxin production.

Amino Acid Sequence↗

Association Analysis of the Circulating Proteome With Sarcopenia-Related Traits Reveals Potential Drug Targets for Sarcopenia.

BACKGROUND: Sarcopenia severely affects the physical health of the elderly. Currently, there is no specific drug available for sarcopenia. This study aims to identify pathogenic proteins and druggable targets for sarcopenia through Mendelian randomization (MR)-based analytical framework. METHODS: A sequential stepwise screening method that includes two-sample MR, Steiger filtering test and colocalization (MRSC) was applied to identify causal proteins associated with sarcopenia-related traits. In the MR analyses, 4372 circulating proteins with valid instrumental variables (IVs) from eight proteomic genome-wide association studies were utilized as exposures, and nine sarcopenia-related traits were utilized as outcomes. IVs were classified into cis-protein quantitative trait loci (pQTLs) and trans-pQTLs based on their positions. We conducted cis-only MRSC analyses and cis&#x2009;+&#x2009;trans MRSC analyses using cis-pQTLs and cis&#x2009;+&#x2009;trans pQTLs as IVs, respectively. Post-MRSC analyses were conducted on the prioritized findings of MRSC, including annotation of protein-altering variants (PAVs), assessment of overlap between pQTLs and expression quantitative trait loci (eQTLs), protein-protein interaction (PPI) analysis, pathway enrichment analysis and annotation of drug targets. Utilizing data from the UK Biobank, we performed an observational study to explore the associations between baseline circulating protein levels and the longitudinal changes in nine sarcopenia-related traits. RESULTS: A total of 181 causal associations for 65 proteins were prioritized by the cis-only MRSC analyses and 227 associations for 91 proteins were prioritized by the cis&#x2009;+&#x2009;trans MRSC analyses. Among the prioritized proteins, the majority of them employed non-PAVs as IVs and most of their cis-pQTLs overlapped with corresponding eQTLs and exhibited consistent directionality, with only one trans-pQTL overlapping with an eQTL. The PPI network of cis-only MRSC-prioritized proteins (p&#x2009;=&#x2009;4.04&#x2009;&#xd7;&#x2009;10-4) and cis&#x2009;+&#x2009;trans MRSC-prioritized proteins (p&#x2009;=&#x2009;8.76&#x2009;&#xd7;&#x2009;10-5) showed significantly more interactions than expected. Reactome, KEGG and GO pathway enrichment analyses for cis-only MRSC-prioritized proteins identified 52, 12 and 79 enriched pathways, respectively (adjusted p&#x2009;<&#x2009;0.05). For proteins identified by cis&#x2009;+&#x2009;trans MRSC analyses, only 15 pathways were enriched through the GO pathway enrichment analyses. In the observational study, 197 circulating proteins were identified to be associated with one or more sarcopenia-related traits (p&#x2009;<&#x2009;0.05/2923). Among them, the significant associations of CTSB (negative association) and ASGR1 (positive association) with sarcopenia-related traits were observed to have consistent directional associations in both MR-based studies and observational studies. Drug target annotations suggested that 52 MRSC-prioritized proteins and 145 biomarkers are drug targets or druggable. CONCLUSIONS: This study identified 89 potential pathogenic proteins and 197 candidate biomarkers for sarcopenia, providing valuable clues for the development of therapeutic drugs for sarcopenia.

Humans↗

Evaluation of UltraSTAR: performance of a collaborative structured data entry system.

The UltraSTAR structured data entry system is now in routine use for reporting ultrasound studies at Brigham and Women's Hospital, having been used for 3722 reports in its first ten months of service. Reports entered through GUI-based forms are uploaded via HL7 to a radiology information system and distributed through a hospital network. UltraSTAR introduces collaborative reporting, in which nonmedical and medical staff collaborate to produce a single report for each patient visit. Performance of UltraSTAR was measured as user satisfaction, data entry time, report completeness, free text annotation rate, and referring-physician satisfaction with reports. Results show high satisfaction with UltraSTAR among radiologists and acceptance of the system among ultrasound technicians. Data entry times averaged 5.3 minutes per report. UltraSTAR reports were slightly more complete than comparable narrative reports. Free text annotations were needed in only 25.2% of all UltraSTAR reports. Referring physicians were neutral to slightly positive toward UltraSTAR's outline-format reports. UltraSTAR is successful at structured data entry despite somewhat long reporting times. Its success can be attributed to efficiencies from collaborative reporting and from integration with existing information systems. UltraSTAR shows that the advantages of structured data entry can outweigh its difficulties even before problems of data entry time and concept representation are solved.

Attitude to Computers↗

CEBS object model for systems biology data, SysBio-OM.

MOTIVATION: To promote a systems biology approach to understanding the biological effects of environmental stressors, the Chemical Effects in Biological Systems (CEBS) knowledge base is being developed to house data from multiple complex data streams in a systems friendly manner that will accommodate extensive querying from users. Unified data representation via a single object model will greatly aid in integrating data storage and management, and facilitate reuse of software to analyze and display data resulting from diverse differential expression or differential profile technologies. Data streams include, but are not limited to, gene expression analysis (transcriptomics), protein expression and protein-protein interaction analysis (proteomics) and changes in low molecular weight metabolite levels (metabolomics). RESULTS: To enable the integration of microarray gene expression, proteomics and metabolomics data in the CEBS system, we designed an object model, Systems Biology Object Model (SysBio-OM). The model is comprehensive and leverages other open source efforts, namely the MicroArray Gene Expression Object Model (MAGE-OM) and the Proteomics Experiment Data Repository (PEDRo) object model. SysBio-OM is designed by extending MAGE-OM to represent protein expression data elements (including those from PEDRo), protein-protein interaction and metabolomics data. SysBio-OM promotes the standardization of data representation and data quality by facilitating the capture of the minimum annotation required for an experiment. Such standardization refines the accuracy of data mining and interpretation. The open source SysBio-OM model, which can be implemented on varied computing platforms is presented here. AVAILABILITY: A universal modeling language depiction of the entire SysBio-OM is available at http://cebs.niehs.nih.gov/SysBioOM/. The Rational Rose object model package is distributed under an open source license that permits unrestricted academic and commercial use and is available at http://cebs.niehs.nih.gov/cebsdownloads. The database and interface are being built to implement the model and will be available for public use at http://cebs.niehs.nih.gov.

Database Management Systems↗

BarleyExpress: a web-based submission tool for enriched microarray database annotations.

UNLABELLED: BarleyExpress is a web-based microarray experiment data submission tool for BarleyBase, a public data resource of Affymetrix GeneChip data for plants. BarleyExpress uses the Plant Ontology vocabularies and enhances the MIAME guidelines to standardize the annotation of microarray gene expression experiments. In addition, BarleyExpress provides explicit support for factorial experiment design and template loading methods to ease the submission process for large experiments. AVAILABILITY: http://barleybase.org SUPPLEMENTARY INFORMATION: BarleyExpress Users Manual.

Database Management Systems↗

Comparative proteome analysis of secretory proteins from pathogenic and nonpathogenic Listeria species.

Extracellular proteins of bacterial pathogens play a crucial role in the infection of the host. Here we present the first comprehensive validation of the secretory subproteome of the Gram positive pathogen Listeria monocytogenes using predictive bioinformatic and experimental proteomic approaches. The previous original signal peptide (SP) prediction (Glaser et al., Science 2001, 294, 849-852) has been greatly improved by an in-depth analysis using seven different bioinformatic tools. Subsequent careful classification of the resulting data gives a probability dependent annotation of 121 putatively secreted proteins of which 45 are novel. Complementary proteomic analysis using both two-dimensional gel electrophoresis/matrix assisted laser desorption/ionization mass spectrometry and high performance liquid chromatography/electrospray ionization-mass spectrometry has identified 105 proteins in the culture supernatant of L. monocytogenes. Among these, we were able to detect all the currently known virulence factors with an SP showing the importance of this subproteome and demonstrating the reliability of the techniques used. The comparison between the L. monocytogenes wildtype and the nonpathogenic species Listeria innocua was performed to reveal proteins probably involved in pathogenicity and/or the adaptation to their respective lifestyles. In addition to the eight known virulence factors, all of which have no orthologous genes in L. innocua, eight additional proteins have been identified that exhibit the typical key feature defining the known listerial virulence factors. Further significant differences between the two species are evident in the group of cell wall and secretory proteins that warrant further study. Our investigation clearly demonstrates that the major difference between the pathogenic and nonpathogenic species, noted in the comparative genome analysis, manifests itself strongest in the secretome.

Bacterial Proteins↗

Organization of the MASP2 locus and its expression profile in mouse and rat.

The mouse, rat, and human MASP2 loci are situated on syntenic chromosome regions and are highly conserved. They comprise the genes for MASP-2/ MAp19, TAR DNA binding protein of 43 kDa, FRAP kinase, CDT6, Polymyositis-Scleroderma 100-kDa autoantigen, spermidine synthase, and TERE which were analyzed by annotation of available gene transcript data and cross-species comparison of available genomic sequences. The human and rat genes for spermidine synthase have an additional intron compared to the mouse gene. The mouse and rat genes for Polymyositis-Scleroderma 100-kDa autoantigen have an additional exon compared to the human gene. We find support for the hypothesis that the MAp19-specific exon within the MASP2 gene may have originated in a transposable element. Blocks of highly conserved intronic sequences were found in the MASP2 gene and the TARDBP gene. The expression of all genes within the MASP2 locus was analyzed in mouse and rat. The restricted expression of MASP-2 and MAp19 mRNA in liver contrasts with the ubiquitous expression of all neighboring genes studied.

Animals↗

Expression of protein elongation factor eEF1A2 predicts favorable outcome in breast cancer.

Breast cancer is the most common malignancy among North American women. The identification of factors that predict outcome is key to individualized disease management and to our understanding of breast oncogenesis. We have analyzed mRNA expression of protein elongation factor eEF1A2 in two independent breast tumor populations of size n = 345 and n = 88, respectively. We find that eEF1A2 mRNA is expressed at a low level in normal breast epithelium but is detectably expressed in approximately 50-60% of primary human breast tumors. We have derived an eEF1A2-specific antibody and measured eEF1A2 protein expression in a sample of 438 primary breast tumors annotated with 20-year survival data. We find that high levels of eEF1A2 protein are detected in 60% of primary breast tumors independent of HER-2 protein expression, tumor size, lymph node status, and estrogen receptor (ER) expression. Importantly, we find that high eEF1A2 is a significant predictor of outcome. Women whose tumor has high eEF1A2 protein expression have an increased probability of 20-year survival compared to those women whose tumor does not express substantial eEF1A2. In addition, eEF1A2 protein expression predicts increased survival probability in those breast cancer patients whose tumor is HER-2 negative or who have lymph node involvement.

Amino Acid Sequence↗

DNA methylation landscape of cerebrospinal fluid cells in multiple sclerosis: an epigenome-wide association study.

BACKGROUND: Multiple sclerosis (MS) is a chronic inflammatory disease of the central nervous system in which DNA methylation may link genetic and environmental risk factors. METHODS: We profiled genome-wide DNA methylation in cerebrospinal fluid (CSF) cells from people with MS (pwMS) and matched controls. Differentially methylated positions (DMPs) and regions (DMRs) were integrated with transcriptomic data, T-cell chromatin annotations, and pathway analyses. Protocadherin gamma (PCDH&#x3b3;) expression was assessed in primary CD4+ T-cell subsets and confirmed by flow cytometry. FINDINGS: We identified 2710 DMPs and 4330 DMRs associating with genes that were enriched in immune signalling, adhesion and migration processes, and were accompanied by corresponding RNA changes. MS-associated methylation changes enriched in the cohesin chromatin-regulation pathway localised to T-cell regulatory regions, and this pathway included multiple protocadherin (PCDH) genes, which displayed consistent methylation and expression changes in CSF cells of pwMS compared to controls. PCDH&#x3b3; cluster gene expression was detected in CD4+ T-cell subsets, and flow cytometry confirmed PCDH&#x3b3; protein expression in peripheral blood T cells. Moreover, co-expression analysis suggests a role of PCDH genes in aryl hydrocarbon receptor (AHR) signalling. Protein-level validation showed fewer PCDH&#x3b3;-positive CD4+ T cells in pwMS and activation-induced PCDH&#x3b3; upregulation after T-cell stimulation. INTERPRETATION: DNA methylation changes in CSF resident cells reflect dysregulated T cell activation and migration in pwMS and suggest involvement of protocadherin molecules in MS pathogenesis. FUNDING: European Research Council, Swedish Research Council, Swedish Brain Foundation, Swedish MS Foundation, Knut and Alice Wallenberg Foundation, European Union and others.

Humans↗

PeroxiBase: a class III plant peroxidase database.

Class III plant peroxidases (EC 1.11.1.7), which are encoded by multigenic families in land plants, are involved in several important physiological and developmental processes. Their varied functions are not yet clearly determined, but their characterization will certainly lead to a better understanding of plant growth, differentiation and interaction with the environment, and hence to many exciting applications. Since there is currently no central database for plant peroxidase sequences and many plant sequences are not deposited in the EMBL/GenBank/DDBJ repository or the UniProt KnowledgeBase, this prevents researchers from easily accessing all peroxidase sequences. Furthermore, gene expression data are poorly covered and annotations are inconsistent. In this rapidly moving field, there is a need for continual updating and correction of the peroxidase superfamily in plants. Moreover, consolidating information about peroxidases will allow for comparison of peroxidases between species and thus significantly help making correlations of function, structure or phylogeny. We report a new database (PeroxiBase) accessible through a web server with specific tools dedicated to facilitate query, classification and submission of peroxidase sequences. Recent developments in the field of plant peroxidase are also mentioned.

Databases, Genetic↗

BioPD: a web-based information center for bioactive peptides.

Bioactive peptide database (BioPD) is a web-based knowledge base that contains more than 1100 protein sequences from human, mouse and rat, which are putative or are known to be bioactive peptides. In addition to peptide sequences and the annotation, the database also contains gene sequences with annotation, protein interaction and disease data related to the peptides. Each entry has as many references as possible to support the information represented. BioPD consists of six parts: PROTEIN, GENE, DISEASE, LINKS, INTERACTION, and REFERENCE. The database is searchable through keyword, gene and protein name, receptor name, etc. The links to PDB, InterPro, Pfam, OMIM, etc. are provided in each entry. Thus BioPD is formed as an information center for the bioactive peptide and serves as a gateway for exploration of bioactive peptides. The database can be accessed at http://biopd.bjmu.edu.cn.

Animals↗

Neuroethology application for the study of human temporal lobe epilepsy: from basic to applied sciences.

The aim of this investigation was to apply neuroethology to the study of human temporal lobe epilepsy (TLE). For this purpose, 42 seizures in 7 patients recorded during video/EEG monitoring (1997-1998) were analyzed by means of a behavioral glossary containing all behaviors. Video recordings were reobserved, and all patients' behaviors were annotated second-by-second. Data were analyzed using Ethomatic software and displayed as flowcharts including frequency, mean duration, and sequential statistic interaction of behavioral items (chi2 > or = 10.827, P<0.001). Flowcharts of (1) a group of seizures from a single patient, (2) the sum of four seizures per patient of two patients with right and five patients with left TLE, and (3) the comparison of left versus right TLE are shown. Well-established data in the literature were confirmed, such as aura (especially epigastric), contralateral lateralization value of dystonia and version, consciousness and language alterations in ictal and postictal periods, mostly with respect to dominant hemisphere involvement, among others. Less well established data such as awakening seizures in TLE patients, lateralization value of facial wiping (ipsilateral to the focus), statistically significant associations between behavioral pairs (dyads), and new behavioral sequences in TLE were also observed. We suggest that neuroethology also has great potential in the study of human epilepsy semiology. This work had an important role in method standardization for human epilepsy, setting the basis for the development of future clinical studies including correlation with other diagnostic methods (EEG, magnetic resonance, and SPECT). The next step will be the comparative study of seizures of patients with left and right TLE, with a greater number of patients, and the development of a digital video library.

Automatism↗