Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 883 records · Page 49Linked to original sources

Mining SARS-CoV protease cleavage data using non-orthogonal decision trees: a novel method for decisive template selection.

MOTIVATION: Although the outbreak of the severe acute respiratory syndrome (SARS) is currently over, it is expected that it will return to attack human beings. A critical challenge to scientists from various disciplines worldwide is to study the specificity of cleavage activity of SARS-related coronavirus (SARS-CoV) and use the knowledge obtained from the study for effective inhibitor design to fight the disease. The most commonly used inductive programming methods for knowledge discovery from data assume that the elements of input patterns are orthogonal to each other. Suppose a sub-sequence is denoted as P2-P1-P1'-P2', the conventional inductive programming method may result in a rule like 'if P1 = Q, then the sub-sequence is cleaved, otherwise non-cleaved'. If the site P1 is not orthogonal to the others (for instance, P2, P1' and P2'), the prediction power of these kind of rules may be limited. Therefore this study is aimed at developing a novel method for constructing non-orthogonal decision trees for mining protease data. RESULT: Eighteen sequences of coronavirus polyprotein were downloaded from NCBI (http://www.ncbi.nlm.nih.gov). Among these sequences, 252 cleavage sites were experimentally determined. These sequences were scanned using a sliding window with size k to generate about 50,000 k-mer sub-sequences (for short, k-mers). The value of k varies from 4 to 12 with a gap of two. The bio-basis function proposed by Thomson et al. is used to transform the k-mers to a high-dimensional numerical space on which an inductive programming method is applied for the purpose of deriving a decision tree for decision-making. The process of this transform is referred to as a bio-mapping. The constructed decision trees select about 10 out of 50,000 k-mers. This small set of selected k-mers is regarded as a set of decisive templates. By doing so, non-orthogonal decision trees are constructed using the selected templates and the prediction accuracy is significantly improved.

Algorithms↗

The risk of gastrointestinal bleed, myocardial infarction, and newly diagnosed hypertension in users of meloxicam, diclofenac, naproxen, and piroxicam.

STUDY OBJECTIVE: To obtain formally quantified data on the relation of meloxicam to newly diagnosed gastrointestinal problems, myocardial infarction, or treated hypertension. DESIGN: Nested case-control study. SETTING: United Kingdom-based General Practice Research Database. PATIENTS: Patients who received prescriptions for meloxicam, diclofenac, naproxen, or piroxicam formed the study population. Cases were people who developed gastrointestinal problems, myocardial infarction, or hypertension. MEASUREMENTS AND MAIN RESULTS: Relative risk estimates for developing the study outcomes were provided for each study nonsteroidal antiinflammatory drug (NSAID), with diclofenac as the reference drug. In no instance was meloxicam associated with an increased risk for a study outcome. CONCLUSION: Compared with the other NSAIDs, meloxicam was not materially associated with any study outcomes. This study provides reassurance to those prescribing this newer class of NSAIDs.

Anti-Inflammatory Agents, Non-Steroidal↗

Identification of an alternative nucleoside triphosphate: 5'-deoxyadenosylcobinamide phosphate nucleotidyltransferase in Methanobacterium thermoautotrophicum delta H.

Computer analysis of the archaeal genome databases failed to identify orthologues of all of the bacterial cobamide biosynthetic enzymes. Of particular interest was the lack of an orthologue of the bifunctional nucleoside triphosphate (NTP):5'-deoxyadenosylcobinamide kinase/GTP:adenosylcobinamide-phosphate guanylyltransferase enzyme (CobU in Salmonella enterica). This paper reports the identification of an archaeal gene encoding a new nucleotidyltransferase, which is proposed to be the nonorthologous replacement of the S. enterica cobU gene. The gene encoding this nucleotidyltransferase was identified using comparative genome analysis of the sequenced archaeal genomes. Orthologues of the gene encoding this activity are limited at present to members of the domain Archaea. The corresponding ORF open reading frame from Methanobacterium thermoautotrophicum Delta H (MTH1152; referred to as cobY) was amplified and cloned, and the CobY protein was expressed and purified from Escherichia coli as a hexahistidine-tagged fusion protein. This enzyme had GTP:adenosylcobinamide-phosphate guanylyltransferase activity but did not have the NTP:AdoCbi kinase activity associated with the CobU enzyme of S. enterica. NTP:adenosylcobinamide kinase activity was not detected in M. thermoautotrophicum Delta H cell extract, suggesting that this organism may not have this activity. The cobY gene complemented a cobU mutant of S. enterica grown under anaerobic conditions where growth of the cell depended on de novo adenosylcobalamin biosynthesis. cobY, however, failed to restore adenosylcobalamin biosynthesis in cobU mutants grown under aerobic conditions where de novo synthesis of this coenzyme was blocked, and growth of the cell depended on the assimilation of exogenous cobinamide. These data strongly support the proposal that the relevant cobinamide intermediates during de novo adenosylcobalamin biosynthesis are adenosylcobinamide-phosphate and adenosylcobinamide-GDP, not adenosylcobinamide. Therefore, NTP:adenosylcobinamide kinase activity is not required for de novo cobamide biosynthesis.

Amino Acid Sequence↗

Stereochemistry of guanidine-metal interactions: implications for L-arginine-metal interactions in protein structure and function.

The geometries of 150 guanidine-metal ion interactions retrieved from crystal structures deposited in the Cambridge Structural Database have been analyzed. Metal ions exhibit a preference for anti coordination stereochemistry in the plane of the unprotonated guanidine group, usually in chelate complexes with a diguanidine moiety, but syn-oriented interactions are occasionally found for single guanidine-metal interactions. Three L-arginine-metal coordination interactions are found in metalloenzyme structures deposited in the Protein Data Bank: biotin synthase from E. coli, His-67 --> Arg human carbonic anhydrase I, and inactivated B. caldovelox arginase complexed with L-arginine. In these proteins, L-arginine-metal coordination adopts syn/out-of-plane and anti/in-plane coordination stereochemistry. The implications of these results for L-arginine-metal interactions in protein structure and function are discussed. Although such interactions are rare, this analysis serves as a useful reference point for the growing interest in enzymes containing L-arginine residues that function as general bases or metal ligands.

Arginine↗

Computer-assisted re-design of spectrin SH3 residue clusters.

We have developed a protein design computer program, called Perla, which performs searches in sequence space to uncover optimal amino acid sequences for desired protein three-dimensional structures. Optimal sequences are localised at the minima of a sequence-structure energy landscape defined using a complex scoring function (an all-atom molecular mechanics force field plus statistical terms including entropy and solvation) measured with respect to a reference state simulating a denatured protein. Sequence choices eventually optimise side chain packing, secondary structure propensities, and hydrogen bonding and electrostatics interactions. Perla was used to re-design clusters of residues of the SH3 domain of alpha-spectrin. Several mutant proteins were produced and characterised. Some of our designed proteins have significantly higher stabilities (stability enhancements about 0.25, 0.70 and 1.0 kcal mol(-1)) than the wild-type protein. These successful protein re-designs, and similar examples found in the literature, establish the quality of the structure-based computational approach to protein design.

Algorithms↗

The clinical spectrum of primary renal vasculitis.

BACKGROUND AND OBJECTIVES: The vasculitides are potentially severe and often difficult to diagnose syndromes. Many forms of vasculitis may involve the kidneys. This review will focus on the clinical and histopathological aspects of renal involvement in the systemic vasculitides. METHODS: We searched the MEDLINE database using as key terms the MeSH terms and textwords for different forms of vasculitis and for renal involvement, creating a database of more than 2200 relevant references. RESULTS: The frequency of renal involvement in vasculitis varies among different syndromes. It is more frequent in Wegener's granulomatosis and microscopic polyarteritis, while it is uncommon to rare in other forms of vasculitis such as Behçet's disease and relapsing polychondritis. The vessels affected include the renal artery in Takayasu arteritis, medium-size renal parenchymal artery in classic polyarteritis nodosa, and glomerular involvement in Wegener's granulomatosis and microscopic polyarteritis. The clinical expression of renal vasculitis depends on the size of the affected vessels and includes renovascular hypertension, isolated nonnephrotic proteinuria, interstitial nephritis, and glomerulonephritis, which can be rapidly progressive. Diagnosis is established by a combination of history, clinical manifestations, laboratory findings (eg, urine sediment, urine protein, antineutrophil cytoplasmic antibodies), imaging techniques (renal angiography, especially when there is a suspicion of medium-to-large vessel disease, and chest radiograph), and finally, renal biopsy. Prognosis varies from unfavorable in the rapidly progressive glomerulonephritis of microscopic polyarteritis, which can lead to renal failure, chronic dialysis, and renal transplantation, to benign, as in the case of Henoch Schonlein purpura, in which the majority of patients recover. CONCLUSIONS: The manifestations and prognosis of renal vasculitis range widely. Renal involvement greatly influences prognosis and dictates the need for early and prompt immunosuppressive therapy. Thus, the clinician should be alert for the timely diagnosis and treatment of renal vasculitis.

Diagnosis, Differential↗

ORFDB: an information resource linking scientific content to a high-quality Open Reading Frame (ORF) collection.

The ORFDB (http://orf.invitrogen.com/) represents an ongoing effort at Invitrogen Corporation to integrate relevant scientific data with an evolving collection of human and mouse Open Reading Frame (ORF) clones (Ultimate ORF Clones). The ORFDB serves as a central data warehouse enabling researchers to search the ORF collection through its web portal ORFBrowser, allowing researchers to find the Ultimate ORF clones by blast, keyword, GenBank accession, gene symbol, clone ID, Unigene ID, LocusLink ID or through functional relationships by browsing the collection via the Gene Ontology (GO) Browser. As of October 2003, the ORFDB contains 6200 human and 2870 mouse Ultimate ORF clones. All Ultimate ORF clones have been fully sequenced with high quality, and are matched to public reference protein sequences. In addition, the cloned ORFs have been extensively annotated across six categories: Gene, ORF, Clone Format, Protein, SNP and Genomic links, with the information assembled in a format termed the ORFCard. The ORFCard represents an information repository that documents the sequence quality, alignment with respect to public protein sequences, and the latest publicly available information associated with each human and mouse gene represented in the collection.

Animals↗

Chemoproteomics-driven drug discovery: addressing high attrition rates.

The advent of multiple high-throughput technologies has brought drug discovery round almost full circle, from pharmacological testing of compounds in vivo to engineered molecular target assays and back to integrated phenotypic screens in cells and organisms. In the past, primary screens to identify new pharmacological agents involved administering compounds to an animal and monitoring a pharmacologic endpoint. For example, antihypertensive agents were identified by dosing spontaneously hypertensive rats with compounds and observing whether their blood pressure dropped. In taking this phenomenological approach, scientists were focused on the final goal, in this example lowering of blood pressure, rather than developing an understanding of the target, or targets, the compounds were impacting. With the evolution of rational target-based approaches, scientists were able to study the direct interaction of compounds with their intended targets, expecting that this would lead to more-selective and safer therapeutics. With the industrialization of screening, referred to as HTS, hundreds of thousands of compounds were screened in robot-driven assays against targets of interest (with this goal in mind). However, an unintentional outcome of the migration from in vivo primary screens to highly target-specific HTS assays was a reduction in biological context caused by the separation of the target from other cellular proteins and processes that might impact its function. Recognition of the potential consequences of this over-simplification drove the modification of HTS processes and equipment to be compatible with cellular assays.

Animals↗

Nuclear receptors, nuclear-receptor factors, and nuclear-receptor-like orphans form a large paralog cluster in Homo sapiens.

We studied a human protein paralog cluster formed by 38 nonredundant sequences taken from the Swiss-Prot database and its supplement, TrEMBL. These sequences include nuclear receptors, nuclear-receptor factors and nuclear-receptor-like orphans. Working separately with both the central cysteine-rich DNA-binding domain and the carboxy-terminal ligand-binding domain, we performed multialignment analyses that included drawings of paralog trees. Our results show that the cluster is highly multibranched, with considerable differences in the amino acid sequence in the ligand-binding domain (LBD), and 17 proximal subbranches which are identifiable and fully coincident when independent trees from both domains are compared. We identified the six recently proposed subfamilies as groups of neighboring clusters in the LBD paralog tree. We found similarities of 80%-100% for the N-terminal transactivation domain among mammalian ortholog receptors, as well as some paralog resemblances within diverse subbranches. Our studies suggest that during the evolutionary process, the three domains were assembled in a modular fashion with a nonshuffled modular fusion of the LBD. We used the EMBL server PredictProtein to make secondary-structure predictions for all 38 LBD subsequences. Amino acid residues in the multialigned homologous domains--taking the beginning of helix H3 of the human retinoic acid receptor-gamma as the initial point of reference--were substituted with H or E, which identify residues predicted to be helical or extended, respectively. The result was a secondary structure multialignment with the surprising feature that the prediction follows a canonical pattern of alignable alpha-helices with some short extended elements in between, despite the fact that a number of subsequences resemble each other by less than 25% in terms of the similarity index. We also identified the presence of a binary patterning in all of the predicted helices that were conserved throughout the 38-sequence sample. Our results fit well with a recently proposed evolutionary model that combines protein secondary structure and amino acid replacement. We propose a new hypothesis for molecular evolution, in which chaperones--acting as an endogenous cellular device for selection--play a crucial role in preserving protein secondary structure.

Amino Acid Sequence↗

PDQ Wizard: automated prioritization and characterization of gene and protein lists using biomedical literature.

SUMMARY: PDQ Wizard automates the process of interrogating biomedical references using large lists of genes, proteins or free text. Using the principle of linkage through co-citation biologists can mine PubMed with these proteins or genes to identify relationships within a biological field of interest. In addition, PDQ Wizard provides novel features to define more specific relationships, highlight key publications describing those activities and relationships, and enhance protein queries. PDQ Wizard also outputs a metric that can be used for prioritization of genes and proteins for further research. AVAILABILITY: PDQ Wizard is freely available from http://www.gti.ed.ac.uk/pdqwizard/.

Abstracting and Indexing↗

Visual analysis of gel-free proteome data.

We present a visual exploration system supporting protein analysis when using gel-free data acquisition methods. The data to be analyzed is obtained by coupling liquid chromatography (LC) with mass spectrometry (MS). LC-MS data have the properties of being nonequidistantly distributed in the time dimension (measured by LC) and being scattered in the mass-to-charge ratio dimension (measured by MS). We describe a hierarchical data representation and visualization method for large LC-MS data. Based on this visualization, we have developed a tool that supports various data analysis steps. Our visual tool provides a global understanding of the data, intuitive detection and classification of experimental errors, and extensions to LC-MS/MS, LC/LC-MS, and LC/LC-MS/MS data analysis. Due to the presence of randomly occurring rare isotopes within the same protein molecule, several intensity peaks may be detected that all refer to the same peptide. We have developed methods to unite such intensity peaks. This deisotoping step is visually documented by our system, such that misclassification can be detected intuitively. For differential protein expression analysis, we compute and visualize the differences in protein amounts between experiments. In order to compute the differential expression, the experimental data need to be registered. For registration, we perform a nonrigid warping step based on landmarks. The landmarks can be assigned automatically using protein identification methods. We evaluate our methods by comparing protein analysis with and without our interactive visualization-based exploration tool.

Chromatography, Gel↗

Improvement in reliability of probabilistic test of significant differences in GeneChip experiments.

A probabilistic test (FUMI theory) for GeneChip experiments has been proposed for selecting the genes which show significant differences in the gene expression levels between a single pair of treatment and control. This paper describes that the reliability of the judgment by the FUMI theory can be enhanced, when the selected genes are referred to biomolecular-functional networks of a commercial database. The genes judged as being differently expressed are grouped into a cluster in the biomolecular networks. It is also demonstrated that false positive genes have a trend in the networks to be isolated from each other, and also away from the clustered genes, since the false positive genes are randomly selected.

Algorithms↗

Inhomogeneous molecular density: reference packing densities and distribution of cavities within proteins.

MOTIVATION: There is no consensus in the literature about how the deepest portions of protein structures are packed. Using an improved Voronoi procedure, we calculate reference packing densities for different regions in the protein interior. Furthermore, we want to clarify where cavities are located. RESULTS: Sets of reference packing densities are provided for regions in proteins that differ in their distance to the surface and to internal cavities, supplementing previous data. Packing in the protein interior is tight but generally inhomogeneous. There are about 4.4 cavities per 100 amino acids in protein structures, they occur in all regions, most frequently in a depth of 2.5-3.6 A underneath the Connolly surface. However, the deepest protein regions have a lower mean packing density than circumjacent regions, because more contacts to cavities occur in the core. AVAILABILITY/SUPPLEMENTARY INFORMATION: Calculation software and detailed packing data are available on request.

Amino Acid Sequence↗

Genome-related datasets within the E. coli Genetic Stock Center database.

The contents of the E. coli Genetic Stock Center database and the availability in electronic form of the subset of information most relevant to sequence databases are described. The database uses the long-standing Stock Center records (developed and curated by Dr B.J.Bachmann) in describing genotypes of mutant derivatives of E.coli K-12 in terms of alleles, structural mutations, mating type, and plasmids as well as the derivation, names and originators of the strain, and references. The database includes descriptions of mutations, mutation properties, genes, gene properties, and gene products, with EC number identifiers for enzymes. Sequence information is not included, but entries refer to sequence database accession numbers for sequenced regions. A gene is described as a subtype of a more general category of chromosome interval called Site. Since sites are used to describe any chromosomal interval, mapping information is associated with sites. Alleles are described as mutations of those sites and they are not primary map objects, but inherit map position information from the corresponding site description. The database design is intended to preserve richness of detail where it is known and uncertainty of measurements or information as it occurs in order to represent the stock center records as accurately as possible.

Bacterial Proteins↗

The evolutionary origin of the protein-translocating channel of chloroplastic envelope membranes: identification of a cyanobacterial homolog.

The known envelope membrane proteins of the chloroplastic protein import apparatus lack sequence similarity to proteins of other eukaryotic or prokaryotic protein transport systems. However, we detected a putative homolog of the gene encoding Toc75, the protein-translocating channel from the outer envelope membrane of pea chloroplasts, in the genome of the cyanobacterium Synechocystis sp. PCC 6803. We investigated whether the low sequence identity of 21% reflects a structural and functional relationship between the two proteins. We provide evidence that the cyanobacterial protein is also localized in the outer membrane. From this information and the similarity of the predicted secondary structures, we conclude that Toc75 and the cyanobacterial protein, referred to as SynToc75, are structural homologs. synToc75 is essential, as homozygous null mutants were not recovered after directed mutagenesis. Sequence analysis indicates that SynToc75 belongs to a family of outer membrane proteins from Gram-negative bacteria whose function is not yet known. However, we demonstrate that these proteins are related to a specific group of prokaryotic secretion channels that transfer virulence factors, such as hemolysins and adhesins, across the outer membrane.

Amino Acid Sequence↗

Generation of a database containing discordant intron positions in eukaryotic genes (MIDB).

MOTIVATION: Intron sliding is the relocation of intron-exon boundaries over short distances and is often also referred to as intron slippage or intron migration or intron drift. We have generated a database containing discordant intron positions in homologous genes (MIDB--Mismatched Intron DataBase). Discordant intron positions are those that are either closely located in homologous genes (within a window of 10 nucleotides) or an intron position that is present in one gene but not in any of its homologs. The MIDB database aims at systematically collecting information about mismatched introns in the genes from GenBank and organizing it into a form useful for understanding the genomics and dynamics of introns thereby helping understand the evolution of genes. RESULTS: Intron displacement or sliding is critically important for explaining the present distribution of introns among orthologous and paralogous genes. MIDB allows examining of intron movements and allows mapping of intron positions from homologous proteins onto a single sequence. The database is of potential use for molecular biologists in general and for researchers who are interested in gene evolution and eukaryotic gene structure. Partial analysis of this database allowed us to identify a few putative cases of intron sliding. AVAILABILITY: http://intron.bic.nus.edu.sg/midb/midb.html

Amino Acid Sequence↗

LuxS controls bacteriocin production in Streptococcus mutans through a novel regulatory component.

The oral pathogen Streptococcus mutans employs a variety of mechanisms to maintain a competitive advantage over many other oral bacteria which occupy the same ecological niche. Production of the bacteriocin, mutacin I, is one such mechanism. However, little is known about the regulatory mechanisms associated with mutacin I production. Previous work has demonstrated that the production of mutacin I greatly increased with cell density. In this study, we found that high cell density also triggered high level mutacin I gene transcription. However, this response was abolished upon deletion of luxS. Further analysis using real-time reverse transcription polymerase chain reaction (RT-PCR) demonstrated that in the luxS mutant transcription of both the mutacin I structural gene mutA and the mutacin I transcriptional activator mutR was impaired. Through microarray analysis, a putative transcription repressor annotated as Smu1274 in the Los Alamos National Laboratory Oral Pathogens Sequence Database was identified, which was strongly induced in the luxS mutant. Characterization of Smu1274, which we referred to as irvA, suggested that it may act as an inducible repressor to suppress mutacin I gene expression. A luxS and irvA double mutant regained the ability to produce mutacin I; whereas a constitutive irvA-producing strain was impaired in mutacin I production. These findings reveal a novel regulatory pathway for mutacin I gene expression, which may provide clues to the regulatory mechanisms of other cellular functions regulated by luxS in S. mutans.

Bacterial Proteins↗

Ranking the whole MEDLINE database according to a large training set using text indexing.

BACKGROUND: The MEDLINE database contains over 12 million references to scientific literature, with about 3/4 of recent articles including an abstract of the publication. Retrieval of entries using queries with keywords is useful for human users that need to obtain small selections. However, particular analyses of the literature or database developments may need the complete ranking of all the references in the MEDLINE database as to their relevance to a topic of interest. This report describes a method that does this ranking using the differences in word content between MEDLINE entries related to a topic and the whole of MEDLINE, in a computational time appropriate for an article search query engine. RESULTS: We tested the capabilities of our system to retrieve MEDLINE references which are relevant to the subject of stem cells. We took advantage of the existing annotation of references with terms from the MeSH hierarchical vocabulary (Medical Subject Headings, developed at the National Library of Medicine). A training set of 81,416 references was constructed by selecting entries annotated with the MeSH term stem cells or some child in its sub tree. Frequencies of all nouns, verbs, and adjectives in the training set were computed and the ratios of word frequencies in the training set to those in the entire MEDLINE were used to score references. Self-consistency of the algorithm, benchmarked with a test set containing the training set and an equal number of references randomly selected from MEDLINE was better using nouns (79%) than adjectives (73%) or verbs (70%). The evaluation of the system with 6,923 references not used for training, containing 204 articles relevant to stem cells according to a human expert, indicated a recall of 65% for a precision of 65%. CONCLUSION: This strategy appears to be useful for predicting the relevance of MEDLINE references to a given concept. The method is simple and can be used with any user-defined training set. Choice of the part of speech of the words used for classification has important effects on performance. Lists of words, scripts, and additional information are available from the web address http://www.ogic.ca/projects/ks2004/.

Abstracting and Indexing↗