Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

Automatic extraction of mutations from Medline and cross-validation with OMIM.

Mutations help us to understand the molecular origins of diseases. Researchers, therefore, both publish and seek disease-relevant mutations in public databases and in scientific literature, e.g. Medline. The retrieval tends to be time-consuming and incomplete. Automated screening of the literature is more efficient. We developed extraction methods (called MEMA) that scan Medline abstracts for mutations. MEMA identified 24,351 singleton mutations in conjunction with a HUGO gene name out of 16,728 abstracts. From a sample of 100 abstracts we estimated the recall for the identification of mutation-gene pairs to 35% at a precision of 93%. Recall for the mutation detection alone was >67% with a precision rate of >96%. This shows that our system produces reliable data. The subset consisting of protein sequence mutations (PSMs) from MEMA was compared to the entries in OMIM (20,503 entries versus 6699, respectively). We found 1826 PSM-gene pairs to be in common to both datasets (cross-validated). This is 27% of all PSM-gene pairs in OMIM and 91% of those pairs from OMIM which co-occur in at least one Medline abstract. We conclude that Medline covers a large portion of the mutations known to OMIM. Another large portion could be artificially produced mutations from mutagenesis experiments. Access to the database of extracted mutation-gene pairs is available through the web pages of the EBI (refer to http://www.ebi. ac.uk/rebholz/index.html).

Animals↗

LISTA, LISTA-HOP and LISTA-HON: a comprehensive compilation of protein encoding sequences and its associated homology databases from the yeast Saccharomyces.

We continued our effort to make a comprehensive database (LISTA) for the yeast Saccharomyces cerevisiae. As in previous editions the genetic names are consistently associated to each sequence with a known and confirmed ORF. If necessary, synonyms are given in the case of allelic duplicated sequences. Although the first publication of a sequence gives-according to our rules-the genetic name of a gene, in some instances more commonly used names are given to avoid nomenclature problems and the use of ancient designations which are no longer used. In these cases the old designation is given as synonym. Thus sequences can be found either by the name or by synonyms given in LISTA. Each entry contains the genetic name, the mnemonic from the EMBL data bank, the codon bias, reference of the publication of the sequence, Chromosomal location as far as known, SWISSPROT and EMBL accession numbers. New entries will also contain the name from the systematic sequencing efforts. Since the release of LISTA4.1 we update the database continuously. To obtain more information on the included sequences, each entry has been screened against non-redundant nucleotide and protein data bank collections resulting in LISTA-HON and LISTA-HOP. This release includes reports from full Smith and Watermann peptide-level searches against a non-redundant protein sequence database. The LISTA data base can be linked to the associated data sets or to nucleotide and protein banks by the Sequence Retrieval System (SRS). The database is available by FTP and on World Wide Web.

Amino Acid Sequence↗

The study of metabolic pathways in tumors based on the transcriptome.

DNA microarray technology revolutionized gene-expression analysis in molecular biology to observe patterns of gene expression in genomic scale. We review the biological aspects of genome-wide gene-expression activity in tumors specially focusing on the analysis of enzyme coding genes. First, the methods for analyzing gene-expression data for the study of metabolome in silico are discussed showing SV40T antigen expressing liver tumor data as an example. Next, an application for tumor metabolome analysis utilizing a reference set of gene-expression profiles is shown.

Animals↗

LISTA, a comprehensive compilation of nucleotide sequences encoding proteins from the yeast Saccharomyces.

The amount of nucleotide sequence data is increasing exponentially. We therefore made an effort to make a comprehensive database (LISTA) for the yeast Saccharomyces cerevisiae. Each sequence has been attributed a single genetic name and in the case of allelic duplicated sequences, synonyms are given, if necessary. For the nomenclature we have introduced a standard principle for naming gene sequences based on priority rules. We have also applied a simple method to distinguish duplicated sequences of one and the same gene from non-allelic sequences of duplicated genes. By using these principles we have sorted out a lot of confusion in the literature and databanks. Along with the genetic name, the mnemonic from the EMBL databank, the codon bias, reference of the publication of the sequence and the EMBL accession numbers are included in each entry.

Base Sequence↗

Molecular interaction map of the mammalian cell cycle control and DNA repair systems.

Eventually to understand the integrated function of the cell cycle regulatory network, we must organize the known interactions in the form of a diagram, map, and/or database. A diagram convention was designed capable of unambiguous representation of networks containing multiprotein complexes, protein modifications, and enzymes that are substrates of other enzymes. To facilitate linkage to a database, each molecular species is symbolically represented only once in each diagram. Molecular species can be located on the map by means of indexed grid coordinates. Each interaction is referenced to an annotation list where pertinent information and references can be found. Parts of the network are grouped into functional subsystems. The map shows how multiprotein complexes could assemble and function at gene promoter sites and at sites of DNA damage. It also portrays the richness of connections between the p53-Mdm2 subsystem and other parts of the network.

Animals↗

A founder mutation in presenilin 1 causing early-onset Alzheimer disease in unrelated Caribbean Hispanic families.

CONTEXT: Genetic determinants of Alzheimer disease (AD) have not been comprehensively examined in Caribbean Hispanics, a population in the United States in whom the frequency of AD is higher compared with non-Hispanic whites. OBJECTIVE: To identify variant alleles in genes related to familial early-onset AD among Caribbean Hispanics. DESIGN AND SETTING: Family-based case series conducted in 1998-2001 at an AD research center in New York, NY, and clinics in the Dominican Republic. PATIENTS: Among 206 Caribbean Hispanic families with 2 or more living members with AD who were identified, 19 (9.2%) had at least 1 individual with onset of AD before the age of 55 years. MAIN OUTCOME MEASURE: The entire coding region of the presenilin 1 gene and exons 16 and 17 of the amyloid precursor protein gene were sequenced in probands from the 19 families and their living relatives. RESULTS: A G-to-C nucleotide change resulting in a glycine-alanine amino acid substitution at codon 206 (Gly206Ala) in exon 7 of presenilin 1 was observed in 23 individuals from 8 (42%) of the 19 families. A Caribbean Hispanic individual with the Gly206Ala mutation and early-onset familial disease was also found by sequencing the corresponding genes of 319 unrelated individuals in New York City. The Gly206Ala mutation was not found in public genetic databases but was reported in 5 individuals from 4 Hispanic families with AD referred for genetic testing. None of the members of these families were related to one another, yet all carriers of the Gly206Ala mutation tested shared a variant allele at 2 nearby microsatellite polymorphisms, indicating a common ancestor. No mutations were found in the amyloid precursor protein gene. CONCLUSIONS: The Gly206Ala mutation was found in 8 of 19 unrelated Caribbean Hispanic families with early-onset familial AD. This genetic change may be a prevalent cause of early-onset familial AD in the Caribbean Hispanic population.

Age of Onset↗

EXProt: a database for proteins with an experimentally verified function.

EXProt is a non-redundant protein database containing a selection of entries from genome annotation projects and public databases, aimed at including only proteins with an experimentally verified function. In EXProt release 2.0 we have collected entries from the Pseudomonas aeruginosa community annotation project (PseudoCAP), the Escherichia coli genome and proteome database (GenProtEC) and the translated coding sequences from the Prokaryotes division of EMBL nucleotide sequence database, which are described as having an experimentally verified function. Each entry in EXProt has a unique ID number and contains information about the species, amino acid sequence, functional annotation and, in most cases, links to references in MEDLINE/PubMed and to the entry in the original database. EXProt is indexed in SRS at CMBI (http://www.cmbi.kun.nl/srs/) and can be searched with BLAST and FASTA through the EXProt web page (http://www.cmbi.kun.nl/EXProt/).

Animals↗

The IMB Jena Image Library of Biological Macromolecules: 2002 update.

The IMB Jena Image Library of Biological Macromolecules (http://www.imb-jena.de/IMAGE.html) is aimed at a better dissemination of information on three-dimensional biopolymer structures with an emphasis on visualization and analysis. It provides access to all structure entries deposited at the Protein Data Bank (PDB) and Nucleic Acid Database (NDB). In addition, basic information on the architecture of biological macromolecules is offered. Recent developments include a site database and an analysis tool that identifies all residues surrounding hetero components or sites according to geometrical criteria. This enables one to search for all structures with a certain pattern of amino acids/nucleotides/water adjacent to hetero components or sites. A new PDB/SWISS-PROT cross-reference database combines information from both PDB and SWISS-PROT, thus providing significantly more cross-references than either PDB or SWISS-PROT. The existing brief descriptions of X-ray, NMR and FTIR methods for structure determination are supplemented by information on circular dichroism.

Animals↗

Proteome analysis on an early transformed human bronchial epithelial cell line, BEP2D, after alpha-particle irradiation.

To probe the mechanism of carcinogenesis of lung cancer at the molecular level and to find potential protein markers involved in the early phase of tumorgenesis, differential proteome analysis on primary passage cell line R15H, and early transformed cell line R15H20 derived from (238)Pu alpha-particle irradiation of human papillomavirus (HPV) 18-immortalized human bronchial epithelial cell line (BEP2D), was carried out using two-dimensional electrophoresis (2-DE) and peptide mass fingerprinting (PMF) with matrix-assisted laser desorption/ionisation-time of flight mass spectrometry. Image analysis and Student's t-test (p < 0.05) showed that three protein spots were only expressed in R15H, intensities of 43 protein spots on the gels were altered between R15H and R15H20. Two of the three spots that were only expressed in R15H were identified as high mobility group protein 1. Two proteins decreased in abundance in R15H20 were identified as maspin precursor, a tumor suppressor and aminoacylase-1. Ornithine aminotransferase and peptidyl-prolyl cis-trans isomerase A that were increased in R15H20, were also identified. Relationships between these differentially expressed proteins and the carcinogenesis mechanism of lung cancer are discussed. The protein expression profile of the R15H cell line was also constructed during the study as a reference map for further comparative proteome analysis of the irradiation induced BEP2D cell line. Of the 90 spots analyzed with PMF in the 2-DE gel of R15H cell line, 50 proteins were identified by searching the nonredundant protein database SWISS-PROT/TrEMBL.

Alpha Particles↗

ProtoBee: hierarchical classification and annotation of the honey bee proteome.

The recently sequenced genome of the honey bee (Apis mellifera) has produced 10,157 predicted protein sequences, calling for a computational effort to extract biological insights from them. We have applied an unsupervised hierarchical protein-clustering method, which was previously used in the ProtoNet system, to nearly 200,000 proteins consisting of the predicted honey bee proteins, the SWISS-PROT protein database, and the complete set of proteins of the mouse (Mus musculus) and the fruit fly (Drosophila melanogaster). The hierarchy produced by this method has been entitled ProtoBee. In ProtoBee, the proteins are hierarchically organized into 18,936 separate tree hierarchies, each representing a protein functional family. By using the mouse and Drosophila complete proteomes as reference, we are able to highlight functional groups of putative gene-loss events, putative novel proteins of unique functionality, and bee-specific paralogs. We have studied some of the ProtoBee findings and suggest their biological relevance. Examples include novel opsin genes and intriguing nuclear matches of mitochondrial genes. The organization of bee sequences into functional clusters suggests a natural way of automatically inferring functional annotation. Following this notion, we were able to assign functional annotation to about 70% of the sequences. ProtoBee is available at http://www.protobee.cs.huji.ac.il.

Animals↗

Comparative evaluation of four urinary tubular dysfunction markers, with special references to the effects of aging and correction for creatinine concentration.

Comparative evaluation was made on alpha(1)-microglobulin (alpha(1)-MG), beta(2)-microglobulin (beta(2)-MG), retinol binding protein (RBP) and N-acetyl-beta-D-glucosaminidase (NAG), as a marker of renal tubular dysfunction after environmental exposure to cadmium (Cd), with special references to the effects of aging and correction for creatinine concentration. For this purpose, a previously established database of 817 never-smoking Japanese women (at the ages of 20 to 74 years) on hematological [hemoglobin, serum ferritin (FE), etc.] and urinary parameters [alpha(1)-MG, beta(2)-MG, creatinine (cr), and a specific gravity] was revisited. For the present analysis, the database was supplemented by the data on RBP and NAG in urine. The exposure of the women to Cd was such that the geometric mean Cd in urine was 1.3 microg/g cr. Among the four tubular dysfunction markers, NAG showed the closest correlation with Cd, followed by alpha(1)-MG and then beta(2)-MG, and RBP was least so although the correlations were all statistically significant. The observed values of the markers gave the best results, whereas correction for a urine specific gravity gave poorer correlation, and it was the worst when correction for creatinine concentration was applied. Age was the most influential confounding factor. The effect of age appeared to be attributable at least in part to the fact that both creatinine and, to a lesser extent, the specific gravity decreased as a function of age. Iron deficiency anemia of sub-clinical degree as observed among the women did not affect any of the four tubular dysfunction markers. In conclusion, NAG and alpha(1)-MG, rather beta(2)-MG or RBP, are more sensitive to detect Cd-induced tubular dysfunction in mass screening. The use of uncorrected observed values of the markers rather than traditional creatinine-corrected values is recommended when comparison covers people of a wide range of ages.

Acetylglucosaminidase↗

Genetic algorithms in molecular recognition and design.

Genetic algorithms provide a novel tool for the investigation of combinatorial optimization problems. A genetic algorithm takes an initial set of possible starting solutions, and iteratively improves them by means of crossover and mutation operators that are related to those involved in Darwinian evolution. This approach is illustrated by reference to applications in molecular modelling, the docking of flexible ligands into protein active sites and de novo ligand design.

Algorithms↗

Retrospective genetic analysis of SAT-1 type foot-and-mouth disease outbreaks in West Africa (1975-1981).

The complete 1D genome region encoding the immunogenic and phylogenetically informative VP1 gene was genetically characterized for 23 South African Territories (SAT)-1 viruses causing foot-and-mouth (FMD) disease outbreaks in the West African region between 1975 and 1981. The results indicate that two independent outbreaks occurred, the first involved two West African countries, namely Niger and Nigeria, whilst the second affected Nigeria alone. In the former epizootic, virus circulation spanned a period of 2 years, whilst in the latter virus was recovered from the field over a 3 year period. Comparison of the West African viruses with SAT-1 viruses from other regions on the continent revealed that the two West African lineages identified in this study are regionally distinct. Furthermore, variation in VP1 gene length was identified in SAT-1 viruses for the first time, further emphasizing the uniqueness of these pathogens in West Africa. This first retrospective analysis in which the molecular epidemiology of SAT-1 viruses in West Africa is reported, provides a useful measure of the regional variation of these viruses and is an essential first step in the establishment of a West African sequence database that will be a useful reference for future outbreak eventualities.

Amino Acid Sequence↗

Applying GIFT, a Gene Interactions Finder in Text, to fly literature.

UNLABELLED: A number of freely available text mining tools have been put together to extract highly reliable Drosophila gene interaction data from text. The system has been tested with The Interactive Fly, showing low recall (27-34%), but very high precision (93-97%). AVAILABILITY: The extracted data and a web interface for submission of texts to GIFT analysis are available at http://gift.cryst.bbk.ac.uk/gift CONTACT: n.domedel_puig@cryst.bbk.ac.uk SUPPLEMENTARY INFORMATION: Additional documentation, such as the dictionaries and the reference sets, are available at the GIFT website.

Artificial Intelligence↗

Integrative missing value estimation for microarray data.

BACKGROUND: Missing value estimation is an important preprocessing step in microarray analysis. Although several methods have been developed to solve this problem, their performance is unsatisfactory for datasets with high rates of missing data, high measurement noise, or limited numbers of samples. In fact, more than 80% of the time-series datasets in Stanford Microarray Database contain less than eight samples. RESULTS: We present the integrative Missing Value Estimation method (iMISS) by incorporating information from multiple reference microarray datasets to improve missing value estimation. For each gene with missing data, we derive a consistent neighbor-gene list by taking reference data sets into consideration. To determine whether the given reference data sets are sufficiently informative for integration, we use a submatrix imputation approach. Our experiments showed that iMISS can significantly and consistently improve the accuracy of the state-of-the-art Local Least Square (LLS) imputation algorithm by up to 15% improvement in our benchmark tests. CONCLUSION: We demonstrated that the order-statistics-based integrative imputation algorithms can achieve significant improvements over the state-of-the-art missing value estimation approaches such as LLS and is especially good for imputing microarray datasets with a limited number of samples, high rates of missing data, or very noisy measurements. With the rapid accumulation of microarray datasets, the performance of our approach can be further improved by incorporating larger and more appropriate reference datasets.

Algorithms↗

Global protein expression pattern of Bradyrhizobium japonicum bacteroids: a prelude to functional proteomics.

As a prelude to using functional proteomics towards understanding the process of symbiotic nitrogen fixation between the legume soybean and the soil bacteria Bradyrhizobium japonicum, we examined the total protein expression pattern of the nodule bacteria, often referred to as bacteroids. A partial proteome map was constructed by separating the total bacteroid proteins using high-resolution 2-DE. Of the several hundred protein spots analyzed using PMF, 180 spots were tentatively identified by searching the available database for B. japonicum, (http://www.kazusa.or.jp/index.html). The data showed that the bacteroid expressed a dominant and elaborate protein network for nitrogen and carbon metabolism, which is closely dependent on the plant supplied metabolites, and seems aptly supported by a selective group of bacteroid transporter proteins. However, they seem to lack a defined fatty acid and nucleic acid metabolism. Interestingly, the proteins related to protein synthesis, scaffolding and degradation were among the most predominant spots of the bacteroid proteome. In addition, several proteins, which showed fairly good expression, were identified to be involved with cellular detoxification, stress regulation and signaling communication components. This preliminary proteomic data matches very well with several biochemical and genetic reports, and clearly shows the inter-connection between several metabolic pathways that meet the needs of the bacteroid. It is expected that in the future this will allow us to develop testable hypotheses about the roles of several of these proteins in context to the metabolic pathway connections and metabolite fluxes.

Amino Acids↗

Normal hematology, serology, and serum protein electrophoresis values in fetal Yucatan miniature swine.

We are currently developing fetal models of congenital heart disease in Yucatan miniature swine for pharmacologic, diagnostic, and interventional methods used to treat cardiac arrhythmias and ventricular septal defect. Fifty-four fetuses from 12 pregnant sows were included in this study. Eleven were fetuses between 76 and 88 days of gestation (early gestation fetuses). A second population of 43 fetuses were between 96 and 110 days of gestation (late gestation fetuses). Erythrocyte, leukocyte, serum electrolyte, enzyme, lipid, carbohydrate, and metabolite values were measured. Complete serum protein profiles were also obtained by electrophoresis. Significant differences could be shown between the sows and fetuses and between the early and late gestation fetuses in all of the categories studied, though not for every parameter. This study provides a large normal database for development of Yucatan miniature swine as an animal model in the rapidly expanding field of fetal medicine.

Animals↗