Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Sensitivity of molecular docking to induced fit effects in influenza virus neuraminidase.

Many proteins undergo small side chain or even backbone movements on binding of different ligands into the same protein structure. This is known as induced fit and is potentially problematic for virtual screening of databases against protein targets. In this report we investigate the limits of the rigid protein approximation used by the docking program, GOLD, through cross-docking using protein structures of influenza neuraminidase. Neuraminidase is known to exhibit small but significant induced fit effects on ligand binding. Some neuraminidase crystal structures caused concern due to the bound ligand conformation and GOLD performed poorly on these complexes. A 'clean' set, which contained unique, unambiguous complexes, was defined. For this set, the lowest energy structure was correctly docked (i.e. RMSD < 1.5 A away from the crystal reference structure) in 84% of proteins, and the most promiscuous protein (1mwe) was able to dock all 15 ligands accurately including those that normally required an induced fit movement. This is considerably better than the 70% success rate seen with GOLD against general validation sets. Inclusion of specific water molecules involved in water-mediated hydrogen bonds did not significantly improve the docking performance for ligands that formed water-mediated contacts but it did prevent docking of ligands that displaced these waters. Our data supports the use of a single protein structure for virtual screening with GOLD in some applications involving induced fit effects, although care must be taken to identify the protein structure that performs best against a wide variety of ligands. The performance of GOLD was significantly better than the GOLD implementation of ChemScore and the reasons for this are discussed. Overall, GOLD has shown itself to be an extremely good, robust docking program for this system.

Algorithms↗

A new human hypervariable locus (K29) maps to the q37.3 region of chromosome 2 and reveals a fingerprint.

A human genomic library was screened with a 30-base oligomer corresponding to the 5' end of the human calretinin cDNA. A clone that contains a minisatellite composed of 21 imperfect repeats of a 37-bp sequence was isolated. The consensus (GAGGGAGGAACTGGGACGCGTGCATGTTTGCATTCTC) incidentally shares 14 consecutive matches with the oligomer used as a probe, and it was shown that the clone did not belong to the calretinin locus. The minisatellite, named K29, was used as a probe on Southern blots at high stringency. After HaeIII, MboI, or HinfI digestion, it detected a single hypervariable locus, with 65% heterozygosity among Caucasian individuals. The probe used at low stringency revealed a fingerprint, with an average of four bands in addition to the locus-specific pattern. Mendelian inheritance was assessed on pedigrees. The K29 minisatellite was mapped by in situ hybridization to the very end of the long arm of chromosome 2 (2q37.3 band), at close proximity of the Fra2J locus, and is referred to as the D2S88 locus in the genome database.

Base Sequence↗

Functional discovery via a compendium of expression profiles.

Ascertaining the impact of uncharacterized perturbations on the cell is a fundamental problem in biology. Here, we describe how a single assay can be used to monitor hundreds of different cellular functions simultaneously. We constructed a reference database or "compendium" of expression profiles corresponding to 300 diverse mutations and chemical treatments in S. cerevisiae, and we show that the cellular pathways affected can be determined by pattern matching, even among very subtle profiles. The utility of this approach is validated by examining profiles caused by deletions of uncharacterized genes: we identify and experimentally confirm that eight uncharacterized open reading frames encode proteins required for sterol metabolism, cell wall function, mitochondrial respiration, or protein synthesis. We also show that the compendium can be used to characterize pharmacological perturbations by identifying a novel target of the commonly used drug dyclonine.

Cell Wall↗

TaxMan: a taxonomic database manager.

BACKGROUND: Phylogenetic analysis of large, multiple-gene datasets, assembled from public sequence databases, is rapidly becoming a popular way to approach difficult phylogenetic problems. Supermatrices (concatenated multiple sequence alignments of multiple genes) can yield more phylogenetic signal than individual genes. However, manually assembling such datasets for a large taxonomic group is time-consuming and error-prone. Additionally, sequence curation, alignment and assessment of the results of phylogenetic analysis are made particularly difficult by the potential for a given gene in a given species to be unrepresented, or to be represented by multiple or partial sequences. We have developed a software package, TaxMan, that largely automates the processes of sequence acquisition, consensus building, alignment and taxon selection to facilitate this type of phylogenetic study. RESULTS: TaxMan uses freely available tools to allow rapid assembly, storage and analysis of large, aligned DNA and protein sequence datasets for user-defined sets of species and genes. The user provides GenBank format files and a list of gene names and synonyms for the loci to analyse. Sequences are extracted from the GenBank files on the basis of annotation and sequence similarity. Consensus sequences are built automatically. Alignment is carried out (where possible, at the protein level) and aligned sequences are stored in a database. TaxMan can automatically determine the best subset of taxa to examine phylogeny at a given taxonomic level. By using the stored aligned sequences, large concatenated multiple sequence alignments can be generated rapidly for a subset and output in analysis-ready file formats. Trees resulting from phylogenetic analysis can be stored and compared with a reference taxonomy. CONCLUSION: TaxMan allows rapid automated assembly of a multigene datasets of aligned sequences for large taxonomic groups. By extracting sequences on the basis of both annotation and BLAST similarity, it ensures that all available sequence data can be brought to bear on a phylogenetic problem, but remains fast enough to cope with many thousands of records. By automatically assisting in the selection of the best subset of taxa to address a particular phylogenetic problem, TaxMan greatly speeds up the process of generating multiple sequence alignments for phylogenetic analysis. Our results indicate that an automated phylogenetic workbench can be a useful tool when correctly guided by user knowledge.

Database Management Systems↗

Potential use of a host associated molecular marker in Enterococcus faecium as an index of human fecal pollution.

Several genotypic and phenotypic microbial source tracking (MST) methods have been proposed and utilized to differentiate groups of microorganisms, usually indicator organisms, for the purpose of tracking sources of fecal pollution. Targeting of host-specific microorganisms is one of the approaches currently being tested. These methods are useful as they circumvent the need to isolate individual microorganisms and do not require the establishment of reference databases. Several studies have demonstrated that the presence and distribution of Enterococcus spp. in feces seems to be influenced by the host species. Here, we present a method for detection of genetic sequences in culturable enterococci capable of identifying human sources of fecal pollution in the environment. The human fecal pollution marker designed in this study targets a putative virulence factor, the enterococcal surface protein (esp), in Enterococcus faecium. This gene was detected in 97% of sewage and septic samples but was not detected in any livestock waste lagoons or in bird or animal fecal samples. Epidemiological studies in recreational and groundwaters have shown enterococci to be useful indicators of public health risk for gastroenteritis. By identifying the presence of human fecal pollution, and therefore the possible presence of human enteric pathogens, this marker allows for further resolution of the source of this risk.

DNA, Bacterial↗

HDL particle associated proteins in plasma and cerebrospinal fluid: identification and partial sequencing.

The proteins from plasma HDL particles isolated by immunoaffinity chromatography on anti-apolipoprotein A-1 affinity columns have been analysed and purified by high resolution two dimensional gel electrophoresis. Two of the lipoprotein-associated proteins found in the HDL plasma fraction, previously referred to as NA1 and NA2, have also been found in cerebrospinal fluid. After separation by 2DGE, these two proteins were transferred to PVDF membranes, stained and cut out for N-terminal sequencing. The partial sequences (11 and 13 amino acids) obtained for the two HDL particle associated proteins do not match any of those included in the December 1987 National Biomedical Research Foundation (NBRF) database, and there are no significant sequence similarities.

Amino Acid Sequence↗

AGML Central: web based gel proteomic infrastructure.

SUMMARY: AGML Central is a web-based open-source public infrastructure for dissemination of two-dimensional Gel Electrophoresis (2-DE) proteomics data in AGML format (Annotated Gel Markup Language). It includes a growing collection of converters from proprietary formats such as those produced by PDQUEST (BioRad), PHORETIX 2-D (Nonlinear Dynamics) and Melanie (GenBio SA). The resulting unifying AGML formatted entry, with or without the raw gel images, is optionally stored in a database for future reference. AGML Central was developed to provide a common platform for data dissemination and development of 2-DE data analysis tools. This resource responds to an increasing use of AGML for 2-DE public source data representation which requires automated tools for conversion from proprietary formats. Conversion and short-term storage is made publicly available, permanent storage requires prior registering. A JAVA applet visualizer was developed to visualize the AGML data with cross-reference links. In order to facilitate automated access a SOAP web service is also included in the AGML Central infrastructure. AVAILABILITY: http://bioinformatics.musc.edu/agmlcentral.

Database Management Systems↗

Towards establishing a protein database of Drosophila.

An improved method of high-resolution two-dimensional gel electrophoresis has been used to study the patterns of protein synthesis in wing imaginal discs of late instar larvae of Drosophila melanogaster. A total of one thousand and twenty five labelled polypeptides (787 acidic and 238 basic) have so far been separated and catalogued. For convenience, all these polypeptides have been numbered and their position fixed by its molecular weight and relative mobility. They are indicated on a reference protein map for further studies.

Animals↗

Pulmonary edema after transfusion: how to differentiate transfusion-associated circulatory overload from transfusion-related acute lung injury.

OBJECTIVE: Pulmonary edema is an under-recognized and potentially serious complication of blood transfusion. Distinct mechanisms include adverse immune reactions and circulatory overload. The former is associated with increased pulmonary vascular permeability and is commonly referred to as transfusion-related acute lung injury (TRALI). The latter causes hydrostatic pulmonary edema and is commonly referred to as transfusion-associated circulatory overload (TACO). In this review article we searched the National Library of Medicine PubMed database as well as references of retrieved articles and summarized the methods for differentiating between hydrostatic and permeability pulmonary edema. RESULTS: The clinical and radiologic manifestations of TACO and TRALI are similar. Although echocardiography and B-type natriuretic peptide measurements may aid in the differential diagnosis between hydrostatic and permeability pulmonary edema, invasive techniques such as right heart catheterization and the sampling of alveolar fluid protein are sometimes necessary. The diagnostic differentiation is especially difficult in critically ill patients will multiple comorbidities so that the cause of edema may only be determined post hoc based on the clinical course and response to therapy. Guided by available evidence, we present an algorithm for establishing the pretest probability of TRALI as opposed to TACO. The decision to test donor and recipient blood for immunocompatibility may be made on this basis. CONCLUSIONS: The distinction between hydrostatic (TACO) and permeability (TRALI) pulmonary edema after transfusion is difficult, in part because the two conditions may coexist. Knowledge of strengths and limitations of different diagnostic techniques is necessary before initiation of complex TRALI workup.

Blood Circulation↗

Re-analysis of data and its integration.

To understand a biological process it is clear that a single approach will not be sufficient, just like a single measurement on a protein--such as its expression level--does not describe protein function. Using reference sets of proteins as benchmarks different approaches can be scaled and integrated. Here, we demonstrate the power of data re-analysis and integration by applying it in a case study to data from deletion phenotype screens and mRNA expression profiling.

Computational Biology↗

NCBI's LocusLink and RefSeq.

The NCBI has introduced two new web resources-LocusLink and RefSeq-that facilitate retrieval of gene-based information and provide reference sequence standards. These resources are designed to provide a non-redundant view of current knowledge about human genes, transcripts and proteins. Additional information about these resources is available on the LocusLink web site at http://www.ncbi.nlm.nih.gov/LocusLink/

Database Management Systems↗

Estimating diversity of Indo-Pacific coral reef stomatopods through DNA barcoding of stomatopod larvae.

There is a push to fully document the biodiversity of the world within 25 years. However, the magnitude of this challenge, particularly in marine environments, is not well known. In this study, we apply DNA barcoding to explore the biodiversity of gonodactylid stomatopods (mantis shrimp) in both the Coral Triangle and the Red Sea. Comparison of sequences from 189 unknown stomatopod larvae to 327 known adults representing 67 taxa in the superfamily Gonodactyloidea revealed 22 distinct larval operational taxonomic units (OTUs). In the Western Pacific, 10 larval OTUs were members of the Gonodactylidae and Protosquillidae where success of positive identification was expected to be 96.5%. However, only five OTUs could be identified to species and at least three OTUs represent new species unknown in their adult form. In the Red Sea where the identification rate was expected to be 75% in the Gonodactylidae, none of four larval OTUs could be identified to species; at least two represent new species unknown in their adult forms. Results indicate that the biodiversity in this well-studied group in the Coral Triangle and Red Sea may be underestimated by a minimum of 50% to more than 150%, suggesting a much greater challenge in lesser-studied groups. Although the DNA barcoding methodology was effective, its overall success was limited due to the newly discovered taxonomic limitations of the reference sequence database, highlighting the importance of synergy between molecular geneticists and taxonomists in understanding and documenting our world's biodiversity, both in marine and terrestrial environments.

Animals↗

Genetic variation in the VP7 gene of human rotavirus isolated in Montevideo-Uruguay from 1996-1999.

The gene encoding the protein VP7 that induces the major neutralizing response has been sequenced from 34 human rotaviruses isolated from children with acute diarrhea in Montevideo (Uruguay) over a 4-year period (1996-1999). These sequences were analyzed and compared to representative corresponding sequences available on databases. In most years, serotype G1 was present as the single serotype, except in 1999 when serotypes G1 and G4 were present simultaneously. Two G1 VP7 lineages were identified. Serotype G2 was present in 1997. The G4 isolates are grouped with Argentine strains and emerged during 1998 in a recently defined sublineage. Neither serotype G3 nor the emerging serotype G9 were isolated during the study. Antigenic domains of isolates and of representative reference strains of each serotype were compared. Sequences of strains isolated during the same year, showed a high degree of homology among strains belonging to the same serotype.

Amino Acid Sequence↗

A diverse superfamily of enzymes with ATP-dependent carboxylate-amine/thiol ligase activity.

The recently developed PSI-BLAST method for sequence database search and methods for motif analysis were used to define and expand a superfamily of enzymes with an unusual nucleotide-binding fold, referred to as palmate, or ATP-grasp fold. In addition to D-alanine-D-alanine ligase, glutathione synthetase, biotin carboxylase, and carbamoyl phosphate synthetase, enzymes with known three-dimensional structures, the ATP-grasp domain is predicted in the ribosomal protein S6 modification enzyme (RimK), urea amidolyase, tubulin-tyrosine ligase, and three enzymes of purine biosynthesis. All these enzymes possess ATP-dependent carboxylate-amine ligase activity, and their catalytic mechanisms are likely to include acylphosphate intermediates. The ATP-grasp superfamily also includes succinate-CoA ligase (both ADP-forming and GDP-forming variants), malate-CoA ligase, and ATP-citrate lyase, enzymes with a carboxylate-thiol ligase activity, and several uncharacterized proteins. These findings significantly extend the variety of the substrates of ATP-grasp enzymes and the range of biochemical pathways in which they are involved, and demonstrate the complementarity between structural comparison and powerful methods for sequence analysis.

Adenosine Triphosphate↗

Estimation of the number of alpha-helical and beta-strand segments in proteins using circular dichroism spectroscopy.

A simple approach to estimate the number of alpha-helical and beta-strand segments from protein circular dichroism spectra is described. The alpha-helix and beta-sheet conformations in globular protein structures, assigned by DSSP and STRIDE algorithms, were divided into regular and distorted fractions by considering a certain number of terminal residues in a given alpha-helix or beta-strand segment to be distorted. The resulting secondary structure fractions for 29 reference proteins were used in the analyses of circular dichroism spectra by the SELCON method. From the performance indices of the analyses, we determined that, on an average, four residues per alpha-helix and two residues per beta-strand may be considered distorted in proteins. The number of alpha-helical and beta-strand segments and their average length in a given protein were estimated from the fraction of distorted alpha-helix and beta-strand conformations determined from the analysis of circular dichroism spectra. The statistical test for the reference protein set shows the high reliability of such a classification of protein secondary structure. The method was used to analyze the circular dichroism spectra of four additional proteins and the predicted structural characteristics agree with the crystal structure data.

Animals↗

Predicting fetal chromosome anomalies in the first trimester using pregnancy associated plasma protein-A: a comparison of statistical methods.

The analysis of the clinical efficiency of a biochemical parameter in the prediction of chromosome anomalies is described, using a database of 475 cases including 30 abnormalities. A comparison was made of two different approaches to the statistical analysis: the use of Gaussian frequency distributions and likelihood ratios, and logistic regression. Both methods computed that for a 5% false-positive rate approximately 60% of anomalies are detected on the basis of maternal age and serum PAPP-A. The logistic regression analysis is appropriate where the outcome variable (chromosome anomaly) is binary and the detection rates refer to the original data only. The likelihood ratio method is used to predict the outcome in the general population. The latter method depends on the data or some transformation of the data fitting a known frequency distribution (Gaussian in this case). The precision of the predicted detection rates is limited by the small sample of abnormals (30 cases). Varying the means and standard deviations (to the limits of their 95% confidence intervals) of the fitted log Gaussian distributions resulted in a detection rate varying between 42% and 79% for a 5% false-positive rate. Thus, although the likelihood ratio method is potentially the better method in determining the usefulness of a test in the general population, larger numbers of abnormal cases are required to stabilise the means and standard deviations of the fitted log Gaussian distributions.

Adult↗

Feasibility in the inverse protein folding protocol.

Methods for protein structure (3D)-sequence (1D) compatibility evaluation (threading) have been developed during the past decade. The protocol in which a sequence can recognize its compatible structure in the structural library (i.e., the fold recognition or the forward-folding search) is available for the structure prediction of new proteins. However, the reverse protocol, in which a structure recognizes its homologous sequences among a sequence database, named the inverse-folding search, is a more difficult application. In this study, we have investigated the feasibility of the latter approach. A structural library, composed of about 400 well-resolved structures with mutually dissimilar sequences, was prepared, and 163 of them had remote homologs in the library. We examined whether they could correctly seek their homologs by both forward- and inverse-folding searches. The results showed that the inverse-folding protocol is more effective than the forward-folding protocol, once the reference states of the compatibility functions are appropriately adjusted. This adjustment only slightly affects the ability of the forward-folding search. We noticed that the scoring, in which a given sequence is re-mounted onto a structure according to the 3D-1D alignment determined by the dynamic programming method, is only effective in the forward-folding protocol and not in the inverse-folding protocol. Namely, the inverse-folding search works significantly better with the score given by the 3D-1D alignment per se, rather than that obtained by the re-mounting. The implications of these results are discussed.

Algorithms↗

Identification and biochemical analysis of GRIN1 and GRIN2.

We have identified the novel Galphaz-binding protein, which is referred to as the G-protein-regulated inducer of neurite outgrowth (GRIN1) using the far-western method. GRIN1 is expressed specifically in brain and binds preferentially to the activated form of alpha subunits of Gz, Gi, and Go. Coexpression of GRIN1 and the activated form of Galphao induce neurite outgrowth in Neuro2a cells. We have further identified two human GRIN1 homologs, GRIN2 and GRIN3, in the database. This article shows that GRIN2 can also bind to the GTP-bound form of Galphao. These findings suggest that the GRIN1 family may function as a downstream effector for Galphao to regulate neurite growth.

Adaptor Proteins, Signal Transducing↗