Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 973 records · Page 54Linked to original sources

Bovine enterovirus 2: complete genomic sequence and molecular modelling of a reference strain and a wild-type isolate from endemically infected US cattle.

Bovine enteroviruses are members of the family Picornaviridae, genus Enterovirus. Whilst little is known about their pathogenic potential, they are apparently endemic in some cattle and cattle environments. Only one of the two current serotypes has been sequenced completely. In this report, the entire genome sequences of bovine enterovirus 2 (BEV-2) strain PS87 and a recent isolate from an endemically infected herd in Maryland, USA (Wye3A) are presented. The recent isolate clearly segregated phylogenetically with sequences representing the BEV-2 serotype, as did other isolates from the endemic herd. The Wye3A isolate shared 82 % nucleotide sequence identity with the PS87 strain and 68 % identity with a BEV-1 strain (VG5-27). Comparison of BEV-2 and BEV-1 deduced protein sequences revealed 72-73 % identity and showed that most differences were single amino acid changes or single deletions, with the exception of the VP1 protein, where both BEV-2 sequences were 7 aa shorter than that of BEV-1. Homology modelling of the capsid proteins of BEV-2 against protein database entries for picornaviruses indicated six significant differences among bovine enteroviruses and other members of the family Picornaviridae. Five of these were on the 'rim' of the proposed enterovirus receptor-binding site or 'canyon' (VP1) and one was near the base of the canyon (VP3). Two of these regions varied enough to distinguish BEV-2 from BEV-1 strains. This is the first report and analysis of full-length sequences for BEV-2. Continued analysis of these wild-type strains should yield useful information for genotyping enteroviruses and modelling enterovirus capsid structure.

Animals↗

Multilocus sequence typing for analyses of clonality of Candida albicans strains in Taiwan.

Multilocus sequence typing (MLST) was used to characterize the genetic profiles of 51 Candida albicans isolates collected from 12 hospitals in Taiwan. Among the 51 isolates, 16 were epidemiologically unrelated, 28 were isolates from 11 critically ill, human immunodeficiency virus (HIV)-negative patients, and 7 were long-term serial isolates from 3 HIV-positive patients. Internal regions of seven housekeeping genes were sequenced. A total of 83 polymorphic nucleotide sites were identified. Ten to 20 different genotypes were observed at the different loci, resulting, when combined, in 45 unique genotype combinations or diploid sequence types (DSTs). Thirty (36.1%) of the 83 individual changes were synonymous and 53 (63.9%) were nonsynonymous. Due to the diploid nature of C. albicans, MLST was more discriminatory than the pulsed-field gel electrophoresis-BssHII-restricted fragment method in discriminating epidemiologically related strains. MLST is able to trace the microevolution over time of C. albicans isolates in the same patient. All but one of the DSTs of our Taiwanese strain collections were novel to the internet C. albicans DST database (http://test1.mlst.net/). The DSTs of C. albicans in Taiwan were analyzed together with those of the reference strains and of the strains from the United Kingdom and United States by unweighted-pair group method using average linkages and minimum spanning tree. Our result showed that the DNA type of each isolate was patient specific and associated with ABC type and decade of isolation but not associated with mating type, anatomical source of isolation, hospital origin, or fluconazole resistance patterns.

AIDS-Related Opportunistic Infections↗

Protein Circular Dichroism Data Bank (PCDDB): data bank and website design.

The Protein Circular Dichroism Data Bank (PCDDB) is a new deposition data bank for validated circular dichroism spectra of biomacromolecules. Its aim is to be a resource for the structural biology and bioinformatics communities, providing open access and archiving facilities for circular dichroism and synchrotron radiation circular dichroism spectra. It is named in parallel with the Protein Data Bank (PDB), a long-existing valuable reference data bank for protein crystal and NMR structures. In this article, we discuss the design of the data bank structure and the deposition website located at http://pcddb.cryst.bbk.ac.uk. Our aim is to produce a flexible and comprehensive archive, which enables user-friendly spectral deposition and searching. In the case of a protein whose crystal structure and sequence are known, the PCDDB entry will be linked to the appropriate PDB and sequence data bank files, respectively. It is anticipated that the PCDDB will provide a readily accessible biophysical catalogue of information on folded proteins that may be of value in structural genomics programs, for quality control and archiving in industrial and academic labs, as a resource for programs developing spectroscopic structural analysis methods, and in bioinformatics studies.

Circular Dichroism↗

Preference functions for prediction of membrane-buried helices in integral membrane proteins.

The preference functions method is described for prediction of membrane-buried helices in membrane proteins. Preference for the alpha-helix conformation of amino acid residue in a sequence is a non-linear function of average hydrophobicity of its sequence neighbors. Kyte-Doolittle hydropathy values are used to extract preference functions from a training data set of integral membrane proteins of partially known secondary structure. Preference functions for beta-sheet, turn and undefined conformation are also extracted by including beta-class soluble proteins of known structure in the training data set. Conformational preferences are compared in tested sequence for each residue and predicted secondary structure is associated with the highest preference. This procedure is incorporated in an algorithm that performs accurate prediction of transmembrane helical segments. Correct sequence location and secondary structure of transmembrane segments is predicted for 20 of 21 reference membrane polypeptides with known crystal structure that were not included in the training data set. Comparison with hydrophobicity plots revealed that our preference profiles are more accurate and exhibit higher resolution and less noise. Shorter unstable or movable membrane-buried alpha-helices are also predicted to exist in different membrane proteins with transport function. For instance, in the sequence of voltage-gated ion channels and glutamate receptors, N-terminal parts of known P-segments can be located as characteristic alpha-helix preference peaks. Our e-mail server: predict@drava.etfos.hr, returns a preference profile and secondary structure prediction for a suspected or known membrane protein when its sequence is submitted.

Algorithms↗

Characterization of colorectal-cancer-related cDNA clones obtained by subtractive hybridization screening.

In an attempt to seek out new factors that are related to colorectal carcinogenesis at the molecular level, subtractive hybridization between cDNA of normal mucosal tissues and mRNA of colorectal carcinoma tissues was performed. Subsequent screenings of the cDNA libraries, constructed from normal mucosal tissues, using the "subtractive probes" generated a total of 46 clones that were expressed in normal mucosa but were either expressed at a significantly reduced level or not expressed at all in cancer tissues. Partial nucleotide sequences of all of these cDNA clones were determined, and sequence homology analyses were performed with the Genbank database. Of the 46 cDNA samples, 44 contained substantial sequence homologies with 32 immunoglobulin gene fragments, a helix-loop-helix basic phosphoprotein gene, an acidic ribosomal phosphoprotein P2 gene, a BLR1 gene for Burkitt's lymphoma receptor 1 gene, D5S419 DNA segment containing (C-A) repeats, a glucokinase (GCK) gene, a Na+, K+-ATPase alpha-subunit gene, a histocompatibility system HLA-DR heavy-chain gene, a dystrophic gene, a mucin (MUC2) gene, a mu-glutathione S-transferase gene, a Menkes disease protein gene, and a 40-kDa keratin intermediate filament precursor gene. The remaining two cDNA clones (now registered under GenBank accession numbers U17714 and U20428) showed few (less than 60%) sequence homologies with any known sequences in the GenBank database and, therefore, may represent novel genes whose expression was down-regulated in human colorectal carcinomas. The possible clinical significance of these findings and the involvement of these two genes in the carcinogenesis of colorectal as well as other cancers are being investigated.

Base Sequence↗

The Human Genome Project--an overview.

The human genome sequence will underpin human biology and medicine in the next century, providing a single, essential reference to all genetic information. The international program to determine the complete DNA sequence (3,000 million bases) is well underway. As of January 2000, 50% of the sequence is available in the public domain. A comprehensive working draft is expected this year, and the entire sequence is projected to be finished in 2003. DNA sequencing is carried out on mapped, overlapping bacterial clones of 150-200 kb. The working draft comprises assembled unfinished sequence and is released immediately in the public domain. The draft sequence of each clone is then completed, by closing any remaining gaps and resolving any ambiguities, before the entire sequence is checked, annotated, and submitted to the public databases. The sequence of each clone is finished to an accuracy of >99.99%. The availability of a reference sequence of the genome provides the basis for studying the nature of sequence variation, particularly single nucleotide polymorphisms (SNPs), in human populations. SNP typing is a powerful tool for genetic analysis, and will enable us to uncover the association of loci at specific sites in the genome with many disease traits. SNPs occur at a frequency of approximately 1 SNP/kb throughout the genome when the sequence of any two individuals is compared. Programs to detect and map SNPs in the human genome are underway with the aim of establishing a SNP map of the genome during the next two years. The human genome sequence will provide a complete description of all the genes. Annotation of the sequence with the gene structures is achieved by a combination of computational analysis (predictive and homology-based) and experimental confirmation by cDNA sequencing. Detecting homologies between newly defined gene products and proteins of known function helps to postulate biochemical functions for them, which can then be tested. Establishing the association of specific genes with disease phenotypes by mutation screening, particularly for monogenic disorders, provides further assistance in defining the functions of some gene products, as well as helping to establish the cause of the disease. As our knowledge of gene sequences and sequence variation in populations increases, we will pinpoint more and more of the genes and proteins that are important in common, complex diseases. A more detailed understanding of the function of the human genome will be achieved as we identify sequences that control gene expression. Given the availability of gene sequences, the expression status of genes in particular tissues can be monitored in parallel. By comparing corresponding genomic sequences in different species (for example: man, mouse, chicken, and zebrafish), regions that have been highly conserved during evolution can be identified, many of which reflect conserved functions such as gene regulation. These approaches promise to greatly accelerate our interpretation of the human genome sequence.

Human Genome Project↗

Progress in the definition of a reference human mitochondrial proteome.

Owing to the complexity of higher eukaryotic cells, a complete proteome is likely to be very difficult to achieve. However, advantage can be taken of the cell compartmentalization to build organelle proteomes, which can moreover be viewed as specialized tools to study specifically the biology and "physiology" of the target organelle. Within this frame, we report here the construction of the human mitochondrial proteome, using placenta as the source tissue. Protein identification was carried out mainly by peptide mass fingerprinting. The optimization steps in two-dimensional electrophoresis needed for proteome research are discussed. However, the relative paucity of data concerning mitochondrial proteins is still the major limiting factor in building the corresponding proteome, which should be a useful tool for researchers working on human mitochondria and their deficiencies.

Databases as Topic↗

The Caenorhabditis briggsae genome contains active CbmaT1 and Tcb1 transposons.

The maT clade of transposons is a group of transposable elements intermediate in sequence and predicted protein structure to mariner and Tc transposons, with a distribution thus far limited to a few invertebrate species. We present evidence, based on searches of publicly available databases, that the nematode Caenorhabditis briggsae has several maT-like transposons, which we have designated as CbmaT elements, dispersed throughout its genome. We also describe two additional transposon sequences that probably share their evolutionary history with the CbmaT transposons. One resembles a fold back variant of a CbmaT element, with long (380-bp) inverted terminal repeats (ITRs) that show a high degree (71%) of identity to CbmaT1. The other, which shares only the 26-bp ITR sequences with one of the CbmaT variants, is present in eight nearly identical copies, but does not have a transposase gene and may therefore be cross mobilised by a CbmaT transposase. Using PCR-based mobility assays, we show that CbmaT1 transposons are capable of excising from the C. briggsae genome. CbmaT1 excised approximately 500 times less frequently than Tcb1 in the reference strain AF16, but both CbmaT1 and Tcb1 excised at extremely high frequencies in the HK105 strain. The HK105 strain also exhibited a high frequency of spontaneous induction of unc-22 mutants, suggesting that it may be a mutator strain of C. briggsae.

Amino Acid Sequence↗

Modeling kinase-substrate specificity: implication of the distance between substrate nucleophilic oxygen and attacked phosphorus of ATP analog on binding affinity.

Molecular dynamics simulations were performed on modeled kinase-substrate complexes in an attempt to establish a relationship between structural features and binding ability of the complexes. We found that the monitored distance between substrate nucleophilic oxygen (OG) and attacked phosphorus (PG) of ATP analog correlated closely with the binding affinity. With reference to 3.3 A, the van der Waals sum of oxygen and phosphorus, the calculated distances of good substrates were close to it whereas those of poor substrates were far apart from it. Therefore, it is reasonable to consider the OG-PG distance as a potential criterion to prefigure the kinase-substrate binding specificity and the simple computational techniques may work as an easy approach to distinguish good substrates from weak or poor substrates.

Adenosine Triphosphate↗

Community nephrology: audit of screening for renal insufficiency in a high risk population.

BACKGROUND: The rate of acceptance onto dialysis programmes has doubled in the past 10 years and is steadily increasing. Early detection and treatment of renal failure slows the rate of progression. Is it feasible to screen for patients who are at increased risk of developing renal failure? We have audited primary care records of patients aged 50-75 years who have either hypertension or diabetes, and are therefore considered to be at high risk of developing renal insufficiency. Our aim was to see whether patients had had their blood pressure measured and urine tested for protein within 12 months, and plasma creatinine measured within 24 months. METHODS: This was a retrospective study of case notes and computer records in 12 general practices from inner and greater London. A total of 16,855 patients were aged 50-75 years. From this age group, 2693 (15.5%) patients were identified as being either hypertensive or diabetic, or both. RESULTS: Of the 2561 records audited, 1359 (53.1%) contained a plasma creatinine measured within 24 months, and 11% of these (150) had a value > 125 micromol/l. This equates to a prevalence of renal insufficiency of > 110,000 patients per million in this group. Forty two patients (28%) had been referred to a nephrologist. Of records audited, 73% contained a blood pressure measurement and 29% contained a test for proteinuria within 12 months. CONCLUSIONS: There is a high prevalence of chronic renal insufficiency in hypertensive and diabetic patients. It is feasible to detect renal insufficiency at a primary care level, but an effective system will require computerized databases that code for age, ethnicity, measurement of blood pressure and renal function, as well as diagnoses.

Aged↗

Two-dimensional gel proteome reference map of blood monocytes.

BACKGROUND: Blood monocytes play a central role in regulating host inflammatory processes through chemotaxis, phagocytosis, and cytokine production. However, the molecular details underlying these diverse functions are not completely understood. Understanding the proteomes of blood monocytes will provide new insights into their biological role in health and diseases. RESULTS: In this study, monocytes were isolated from five healthy donors. Whole monocyte lysates from each donor were then analyzed by 2D gel electrophoresis, and proteins were detected using Sypro Ruby fluorescence and then examined for phosphoproteomes using ProQ phospho-protein fluorescence dye. Between 1525 and 1769 protein spots on each 2D gel were matched, analyzed, and quantified. Abundant protein spots were then subjected to analysis by mass spectrometry. This report describes the protein identities of 231 monocyte protein spots, which represent 164 distinct proteins and their respective isoforms or subunits. Some of these proteins had not been previously characterized at the protein level in monocytes. Among the 231 protein spots, 19 proteins revealed distinct modification by protein phosphorylation. CONCLUSION: The results of this study offer the most detailed monocyte proteomic database to date and provide new perspectives into the study of monocyte biology.

Journal Article↗

The peptaibol database: a sequence and structure resource.

The peptaibols are a large family of membrane-active peptides with considerable sequence homology, but with different biological properties and three-dimensional structures. They constitute a rich resource of naturally occurring 'mutants' which are potentially valuable for structure/function studies of ion channels. A searchable on-line database of sequences and structures of the peptaibols has been created at http://www.cryst.bbk.ac.uk/peptaibol, as a resource for the biological and structural community. In this paper, the contents and organization of the website are discussed as well as procedures for submission of new entries to the database. At present, more than 300 peptaibol sequences are stored in the database. Each sequence entry contains its full literature reference and information about its biological source. Tools are provided for searching for specific peptaibol sequences or groupings of sequences, and for locating peptaibols containing specified sequence motifs. In addition the website acts as a database for structural information. The coordinates of all currently available peptaibol x-ray and NMR structures are included and complemented, where appropriate. with molecular graphics illustrations. These include figures of model channel structures and comparisons between different peptaibol structures. The peptaibol database thus provides a tool for ready access to information and a means of investigating the sequences and structures of this class of polypeptides.

Amino Acid Sequence↗

Identification of TINO: a new evolutionarily conserved BCL-2 AU-rich element RNA-binding protein.

Modulation of mRNA stability by regulatory cis-acting AU-rich elements (AREs) and ARE-binding proteins is an important posttranscriptional mechanism of gene expression control. We previously demonstrated that the 3'-untranslated region of BCL-2 mRNA contains an ARE that accounts for rapid BCL-2 down-regulation in response to apoptotic stimuli. We also demonstrated that the BCL-2 ARE core interacts with a number of ARE-binding proteins, one of which is AU-rich factor 1/heterogeneous nuclear ribonucleoprotein D, known for its interaction with mRNA elements of others genes. In an attempt to search for other BCL-2 mRNA-binding proteins, we used the yeast RNA three-hybrid system assay and identified a novel human protein that interacts with BCL-2 ARE. We refer to it as TINO. The predicted protein sequence of TINO reveals two amino-terminal heterogeneous nuclear ribonucleoprotein K homology motifs for nucleic acid binding and a carboxyl-terminal RING domain, endowed with a putative E3 ubiquitin-protein ligase activity. In addition the novel protein is evolutionarily conserved; the two following orthologous proteins have been identified with protein-protein BLAST: posterior end mark-3 (PEM-3) of Ciona savignyi and muscle excess protein-3 (MEX-3) of Caenorhabditis elegans. Upon binding, TINO destabilizes a chimeric reporter construct containing the BCL-2 ARE sequence, revealing a negative regulatory action on BCL-2 gene expression at the posttranscriptional level.

3' Untranslated Regions↗

Comparative analysis of mouse NotI linking clones with mouse and human genomic sequences and transcripts.

NotI cleavage sites are frequently associated with CpG islands that identify the 5' regulatory sites of functional genes in the genome. Therefore we analyzed a sample of 22 NotI linking clones prepared from mouse brain DNA, to determine whether these mouse NotI site associated clones could be used for comparative analysis of mouse and human genomes by cross-reaction with both mouse and human genomic DNA and RNA in Southern and Northern hybridization. We further examined whether we could establish the identity of these clones with known genes by comparing the nucleotide sequences surrounding the NotI site with the GenBank database. We observed that 70% of the clones cross-hybridized with human DNA and that 4 of 11 tested clones (36%) detected a transcript in human HeLa cells RNA whereas 73% clones (8/11) detected transcripts in mouse RNAs from one or more organs. Single pass sequence analysis was successful on 16 of 19 clones. The GC content in these sequence was very high (48.8% to 73.8%) suggesting that 12 of 16 sequenced clones contained a CpG island. Three out of 19 clones showed significant similarity with previously analyzed mouse gene sequences in GenBank, including the mouse rRNA gene family, cathepsin and the scip POU-domain genes. In addition, two sequences showed significant similarity to the human and rabbit protein phosphatase 2A-beta subunit and the human transforming growth factor-beta. Thus, 5 of 16 clones showed homology with identified genes. These results and the recent work of using RLGS methods for genetic mapping indicate that NotI linking clones can be used to efficiently cross reference a comparative analysis of the mouse and human genomic maps.

Animals↗

Protein NMR recall, precision, and F-measure scores (RPF scores): structure quality assessment measures based on information retrieval statistics.

One of the most important challenges in modern protein NMR is the development of fast and sensitive structure quality assessment measures that can be used to evaluate the "goodness-of-fit" of the 3D structure with NOESY data, to indicate the correctness of the fold and accuracy of the resulting structure. Quality assessment is especially critical for automated NOESY interpretation and structure determination approaches. This paper describes new NMR quality assessment scores, including Recall, Precision, and F-measure scores (referred to here are "NMR RPF" scores), which quickly provide global measures of the goodness-of-fit of the 3D structures with NOESY peak lists using methods from information retrieval statistics. The sensitivity of the F-measure is improved using a scaled Fold Discriminating Power (DP) score. These statistical RPF scores are quite rapid to compute since NOE assignments and complete relaxation matrix calculations are not required. A graphical method for site-specific assessment of structure quality based on the Precision statistic is also described. These statistical measures are demonstrated to be valuable for assessing protein NMR structure accuracy. Their relationships to other proposed NMR "R-factors" and structure quality assessment scores are also discussed.

Databases, Protein↗

Transforming a set of biological flat file libraries to a fast access network.

SRS (Sequence Retrieval System), an indexing system for flat file libraries, provides fast access to individual library entries via retrieval by keywords from various data fields. SRS is now also able to build indices using cross-references that most libraries provide. Fifteen libraries of DNA and protein sequences and structures have been selected. These libraries interact with at least one other by means of cross-references. Indexing these cross-references allows a complete network of libraries to be built. In the network an entry from one library can be linked in principle to every other library. If two libraries are not directly cross-referenced, the linkage can be made with a succession of single links between neighbouring, cross-referenced libraries. A new operator has been added to the query language of SRS for convenient specification of links amongst complete libraries or entry sets generated by previous queries on particular libraries. All the information in the network can now be used to retrieve an entry in a specific library, e.g. the full information given in amino acid sequence entries from SwissProt can now be used to retrieve related tertiary structure entries from PDB. Furthermore, a search in a single library can be extended to a search in the complete library network, e.g. all entries in all databases pertaining to elastase can be found.

Algorithms↗

Entamoeba histolytica: construction and applications of subgenomic databases.

Knowledge about the influence of environmental stress such as the action of chemotherapeutic agents on gene expression in Entamoeba histolytica is limited. We plan to use oligonucleotide microarray hybridization to approach these questions. As the basis for our array, sequence data from the genome project carried out by the Institute for Genomic Research (TIGR) and the Sanger Institute were used to annotate parts of the parasite genome. Three subgenomic databases containing enzymes, cytoskeleton genes, and stress genes were compiled with the help of the ExPASy proteomics website and the BLAST servers at the two genome project sites. The known sequences from reference species, mostly human and Escherichia coli, were searched against TIGR and Sanger E. histolytica sequence contigs and the homologs were copied into a Microsoft Access database. In a similar way, two additional databases of cytoskeletal genes and stress genes were generated. Metabolic pathways could be assembled from our enzyme database, but sometimes they were incomplete as is the case for the sterol biosynthesis pathway. The raw databases contained a significant number of duplicate entries which were merged to obtain curated non-redundant databases. This procedure revealed that some E. histolytica genes may have several putative functions. Representative examples such as the case of the delta-aminolevulinate synthase/serine palmitoyltransferase are discussed.

5-Aminolevulinate Synthetase↗

Significance of the genetic relationships deduced from partial nucleotide sequencing of infectious bursal disease virus genome segments A or B.

The rapid genomic characterization of infectious bursal disease virus (IBDV) requires determining which partial nucleotide (nt) sequences derived from IBDV segments A or B would produce phylogenetic information as significant as sequencing the whole corresponding segments. Long nt coding sequences of 27 IBDV segments A (aa 20-991) and 21 segments B (aa 7-stop codon) were retrieved from databanks and used to compute reference phylogenetic trees using Neighbor Joining (NJ) and Parsimony (P): clusters appearing in the NJ and P reference trees with a bootstrap value greater than 80% were considered as significant (Whole Segment Clusters, WSC). The sequences were then cut into overlapping regions. These were used to compute phylogenetic trees which were compared with reference ones. Of the partial sequences, the VP2 gene best represented IBDV segment A (10 out of 13 WSC were conserved), and the 5' two thirds of segment B best represented segment B (5 to 6 conserved WSC out of 6). Implementation of the Plato programme finally demonstrated that the region encoding VP2 variable domain (vVP2, segment A) is the only region of IBDV genome with a significantly different evolution rate, which result is consistent with vVP2 being subjected to a high selection pressure.

Databases, Genetic↗