Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

Nucleotide sequence analysis of nucleocapsid protein gene of canine distemper virus isolates in Thailand.

The C-terminal part of the nucleocapsid protein gene of 13 canine distemper virus (CDV) isolates from Thailand, were analyzed. The nucleotide sequences were assigned to two clusters; cluster A exhibited a high degree of homology with the vaccine strain Onderstepoort, 99.10 and 97.61%, respectively, in the two isolates examined. Cluster B appeared closely related to virulent strains registered in the GeneBank database and to the virulent reference strain (A75/17); a total of 11 samples were analyzed, with 94.63-99.10% homology at the same position. The deduced amino acid sequences correlated with the two-nucleotide sequence clusters. However, there was no association among the CDV groups with histories of vaccination, sex, ages, clinical findings and evidence of viral antigen in tissues.

Amino Acid Sequence↗

IDconverter and IDClight: conversion and annotation of gene and protein IDs.

BACKGROUND: Researchers involved in the annotation of large numbers of gene, clone or protein identifiers are usually required to perform a one-by-one conversion for each identifier. When the field of research is one such as microarray experiments, this number may be around 30,000. RESULTS: To help researchers map accession numbers and identifiers among clones, genes, proteins and chromosomal positions, we have designed and developed IDconverter and IDClight. They are two user-friendly, freely available web server applications that also provide additional functional information by mapping the identifiers on to pathways, Gene Ontology terms, and literature references. Both tools are high-throughput oriented and include identifiers for the most common genomic databases. These tools have been compared to other similar tools, showing that they are among the fastest and the most up-to-date. CONCLUSION: These tools provide a fast and intuitive way of enriching the information coming out of high-throughput experiments like microarrays. They can be valuable both to wet-lab researchers and to bioinformaticians.

Algorithms↗

Automated prediction of 15N, 13Calpha, 13Cbeta and 13C' chemical shifts in proteins using a density functional database.

A database of peptide chemical shifts, computed at the density functional level, has been used to develop an algorithm for prediction of 15N and 13C shifts in proteins from their structure; the method is incorporated into a program called SHIFTS (version 4.0). The database was built from the calculated chemical shift patterns of 1335 peptides whose backbone torsion angles are limited to areas of the Ramachandran map around helical and sheet configurations. For each tripeptide in these regions of regular secondary structure (which constitute about 40% of residues in globular proteins) SHIFTS also consults the database for information about sidechain torsion angle effects for the residue of interest and for the preceding residue, and estimates hydrogen bonding effects through an empirical formula that is also based on density functional calculations on peptides. The program optionally searches for alternate side-chain torsion angles that could significantly improve agreement between calculated and observed shifts. The application of the program on 20 proteins shows good consistency with experimental data, with correlation coefficients of 0.92, 0.98, 0.99 and 0.90 and r.m.s. deviations of 1.94, 0.97, 1.05, and 1.08 ppm for 15N, 13Calpha, 13Cbeta and 13C', respectively. Reference shifts fit to protein data are in good agreement with 'random-coil' values derived from experimental measurements on peptides. This prediction algorithm should be helpful in NMR assignment, crystal and solution structure comparison, and structure refinement.

Algorithms↗

Identification of a novel steroid derivative, NSC12983, as a paclitaxel-like tubulin assembly promoter by 3-D virtual screening.

By flexibly docking 9593 compounds of the NCI-3D database against the refined structure of beta-tubulin using DOCK 4.0, a long-forgotten synthetic steroid derivative, NSC12983, has been identified as a microtubule-stabilizing agent. The 32 top scorers includes NSC12983 and the three added references: paclitaxel, docetaxel and IDN5109. That is, 12.5% of the 0.33% top scorers are active. In addition, NSC12983 is active on Mycobacterium tuberculosis in vitro and in vivo, which might be due to its ability to promote the assembly of essential cell division protein.

Antineoplastic Agents, Phytogenic↗

Open-and-shut cases in coiled-coil assembly: alpha-sheets and alpha-cylinders.

The coiled coil is a ubiquitous protein-folding motif. It generally is accepted that coiled coils are characterized by sequence patterns known as heptad repeats. Such patterns direct the formation and assembly of amphipathic alpha-helices, the hydrophobic faces of which interface in a specific manner first proposed by Crick and termed "knobs-into-holes packing". We developed software, SOCKET, to recognize this packing in protein structures. As expected, in a trawl of the protein data bank, we found examples of canonical coiled coils with a single contiguous heptad repeat. In addition, we identified structures with multiple, overlapping heptad repeats. This observation extends Crick's original postulate: Multiple, offset heptad repeats help explain assemblies with more than two helices. Indeed, we have found that the sequence offset of the multiple heptad repeats is related to the coiled-coil oligomer state. Here we focus on one particular sequence motif in which two heptad repeats are offset by two residues. This offset sets up two hydrophobic faces separated by approximately 150 degrees -160 degrees around the alpha-helix. In turn, two different combinations of these faces are possible. Either similar or opposite faces can interface, which leads to open or closed multihelix assemblies. Accordingly, we refer to these two forms as alpha-sheets and alpha-cylinders. We illustrate these structures with our own predictions and by reference to natural variants on these designs that have recently come to light.

Amino Acid Motifs↗

C alpha and C beta carbon-13 chemical shifts in proteins from an empirical database.

We have constructed an extensive database of 13C C alpha and C beta chemical shifts in proteins of solution, for proteins of which a high-resolution crystal structure exists, and for which the crystal structure has been shown to be essentially identical to the solution structure. There is no systematic effect of temperature, reference compound, or pH on reported shifts, but there appear to be differences in reported shifts arising from referencing differences of up to 4.2 ppm. The major factor affecting chemical shifts is the backbone geometry, which causes differences of ca. 4 ppm between typical alpha-helix and beta-sheet geometries for C alpha, and of ca. 2 ppm for C beta. The side-chain dihedral angle chi 1 has an effect of up to 0.5 ppm on the C alpha shift, particularly for amino acids with branched side-chains at C beta. Hydrogen bonding to main-chain atoms has an effect of up to 0.9 ppm, which depends on the main-chain conformation. The sequence of the protein and ring-current shifts from aromatic rings have an insignificant effect (except for residues following proline). There are significant differences between different amino acid types in the backbone geometry dependence; the amino acids can be grouped together into five different groups with different phi, psi shielding surfaces. The overall fit of individual residues to a single non-residue-specific surface, incorporating the effects of hydrogen bonding and chi 1 angle, is 0.96 ppm for both C alpha and C beta. The results from this study are broadly similar to those from ab initio studies, but there are some differences which could merit further attention.

Animals↗

Molecular similarity based on DOCK-generated fingerprints.

An alternative method for defining molecular similarity is presented. By using the docking program DOCK and a reference panel of protein binding sites, fingerprints for a set of molecules have been generated, based on calculated interaction energies. These binding patterns allowed us to calculate matrices of similarity coefficients which subsequently were used for nearest-neighbor searches within the database. Our results indicate that the method is suitable for finding significant similarities of compounds of the same biological activity. Although the overall performance of a traditional 2D similarity method is better in the test systems investigated, our 3D approach can be regarded as complementary since it is able to detect similarities independent of the covalent structure of the compounds. Thus it should be a useful 3D database-searching tool for rational lead discovery.

Amino Acid Sequence↗

The RESID Database of protein structure modifications.

Because the number of post-translational modifications requiring standardized annotation in the PIR-International Protein Sequence Database was large and steadily increasing, a database of protein structure modifications was constructed in 1993 to assist in producing appropriate feature annotations for covalent binding sites, modified sites and cross-links. In 1995 RESID was publicly released as a PIR-International text database distributed on CD-ROM and accessible through the ATLAS program. In 1998 it was made available on the PIR Web site at http://www-nbrf.georgetown.edu/pir/searchdb++ +.html . The RESID Database includes such information as: systematic and frequently observed alternate names; Chemical s Service registry numbers; atomic formulas and weights; enzyme activities; indicators forN-terminal, C-terminal or peptide chain cross-link modifications; keywords; and literature citations with database cross-references. The RESID Database can be used to predict atomic masses for peptides, and is being enhanced to provide molecular structures for graphical presentation on the PIR Web site using widely available molecular viewing programs.

Binding Sites↗

Proteomic analysis of log to stationary growth phase Lactobacillus plantarum cells and a 2-DE database.

Lactobacillus plantarum is part of the natural microbiota of many food fermentations as well as the human gastro-intestinal tract. The cytosolic fraction of the proteome of L. plantarum WCFS1, whose genome has been sequenced, was studied. 2-DE was used to investigate the proteins from the cytosolic fraction isolated from mid- and late-log, early- and late-stationary phase cells to generate reference maps of different growth conditions offering more knowledge of the metabolic behavior of this bacterium. From this fraction, a total of 200 protein spots were identified by MALDI-MS and a proteome production map was constructed to facilitate further studies such as detection of suitable biomarkers for specific growth conditions. More than half (57%) of the identified proteins were predicted to be involved in metabolic pathways of the bacterium. The protein profile changed during the growth of the bacteria such that 29% of the identified proteins involved in anabolic pathways were at least twofold up-regulated throughout the mid- and late-exponential and early-stationary phases. In the late-stationary phase, six proteins involved in stress or with a potential role for survival during starvation were up-regulated significantly.

Bacterial Proteins↗

Characterization of the LARGE family of putative glycosyltransferases associated with dystroglycanopathies.

The Large(myd) mouse has a loss-of-function mutation in the putative glycosyltransferase gene Large. Mutations in the human homolog (LARGE) have been described in a form of congenital muscular dystrophy (MDC1D). Other genes (POMT1, POMGnT1, fukutin, and FKRP) that encode known or putative glycosylation enzymes are also causally associated with human congenital muscular dystrophies. All these diseases are associated with hypoglycosylation of the membrane protein alpha-dystroglycan (alpha-DG) and consequent loss of extracellular ligand binding. Hence, they are termed dystroglycanopathies. A paralogous gene for LARGE (LARGE2 or GYLTL1B) may also have a role in DG glycosylation. Using database interrogation and reverse-transcriptase polymerase chain reaction (RT-PCR), we identified vertebrate orthologs of each of these LARGE genes in many vertebrates, including human, mouse, dog, chicken, zebrafish, and pufferfish. However, within invertebrate genomes, we were able to identify only single homologs. We suggest that vertebrate LARGE orthologs be referred to as LARGE1. RT-PCR, dot-blot, and northern analysis indicated that LARGE2 has a more restricted tissue-expression profile than LARGE1. Using epitope-tagged proteins, we show that both LARGE1 and LARGE2 localize to the Golgi apparatus. The high similarity between the LARGE paralogs suggests that LARGE2 may also act on DG. Overexpression of LARGE2 in mouse C2C12 myoblasts results in increased glycosylation of alpha-DG accompanied by an increase in laminin binding. Thus, there may be functional redundancy between LARGE1 and LARGE2. Consistent with this idea, we show that alpha-DG is still fully glycosylated in kidney (a tissue that expresses a high level of LARGE2 mRNA) of Large(myd) mutant mice.

Alternative Splicing↗

Protein three-dimensional structure determination and sequence-specific assignment of 13C and 15N-separated NOE data. A novel real-space ab initio approach.

The sequence-specific assignment of resonances is considered to be a requirement for the determination of the three-dimensional (3D) structure of a protein in solution by nuclear magnetic resonance methods. The main source of structural information is the nuclear Overhauser effect spectroscopy (NOESY) spectrum, which contains information about spatially close pairs of protons. Currently, various J-correlated spectra must be recorded in order to obtain the sequence-specific assignments necessary to interpret the NOESY spectra. In this work, a novel procedure to determine the 3D structure and the sequence-specific assignments of a protein using only data from 13C and 15N-separated multidimensional NOESY spectra is described. No information from J-correlated spectra is required. The algorithm is called ANSRS (Assignment of NOESY Spectra in Real Space) and is based on an inversion of the traditional strategy. A 3D real-space structure of detected, but unassigned, 1H spins is calculated from the nuclear Overhauser effect (NOE) distance restraints using a dynamical simulated annealing procedure. The sequence-specific assignments are then determined by searching among the 1H spins in the 3D real-space structure for plausible residue assignments. The search uses a Monte Carlo simulated annealing algorithm based on assignment probabilities derived from the 1H, 15N and 13C chemical shifts, various spatial constraints, and the known sequence of the protein. The procedure has been tested on semi-synthetic data sets comprising published experimental chemical shifts and NOE distance restraints derived from the known 3D structures of the two proteins GAL4 (residues 9 to 41) and bovine pancreatic trypsin inhibitor. The ANSRS procedure was able to determine the sequence-specific assignments for more than 95% of the spins, and was fairly robust with respect to missing NOE data. The potential of the ANSRS approach with respect to automated assignment, reduction of the number of NMR spectra required for a structure determination, assignment of homologous and mutant proteins, and the possibility of analysing spectra recorded at high pH is discussed.

Algorithms↗

Metal-ligand geometry relevant to proteins and in proteins: sodium and potassium.

In previous papers [Harding (2001), Acta Cryst. D57, 401-411, and references therein] the geometry of metal-ligand interactions was examined for six metals (Ca, Mg, Mn, Fe, Cu, Zn) using the Protein Data Bank and compared with information from accurately determined structures of relevant small-molecule crystals in the Cambridge Structural Database. Here, the environments of Na(+) and K(+) ions found in protein crystal structures are examined in an equivalent way. Target M(+).O distances are proposed and the agreement with observed distances is summarized. The commonest interactions are with water molecules and the next commonest with main-chain carbonyl O atoms.

Crystallography↗

DICHROWEB, an online server for protein secondary structure analyses from circular dichroism spectroscopic data.

The DICHROWEB web server enables on-line analyses of circular dichroism (CD) spectroscopic data, providing calculated secondary structure content and graphical analyses comparing calculated structures and experimental data. The server is located at http://www.cryst.bbk.ac.uk/cdweb and may be accessed via a password-limited user ID, available upon completion of a registration form. The server facilitates analyses using five popular algorithms and (currently) seven different reference databases by accepting data in a user-friendly manner in a wide range of formats, including those output by both commercial CD instruments and synchrotron radiation-based circular dichroism beamlines, as well as those produced by spectral processing software packages. It produces as output calculated secondary structures, a goodness-of-fit parameter for the analyses, and tabular and graphical displays of experimental, calculated and difference spectra. The web pages associated with the server provide information on CD spectroscopic methods and terms, literature references and aids for interpreting the analysis results.

Algorithms↗

The use of composite crystal-field environments in molecular recognition and the de novo design of protein ligands.

Small molecule crystal data have been retrieved from the Cambridge Crystallographic Database to compile composite crystal-field environments about different functional groups, which also occur in proteins and nucleotides. Their spatial distribution can be used to map-out putative interaction sites, e.g. about amino acid residues oriented towards the binding site of a given protein. Although influenced by packing forces, these composite environments show systematic patterns which reflect preferred interaction geometries of the functional groups under consideration with neighboring groups, e.g. hydrogen bonding partners. Similar but substantially less detailed distributions have been obtained from crystallographically determined ligand/protein complexes, which demonstrate that the properties observed in low-molecular weight structures are representative also for the sought after spatial orientation of interactions between ligands and their receptor proteins. The crystallographically determined binding geometries of three inhibitor/enzyme complexes are compared with the distributions of putative interaction sites predicted from corresponding composite field environments. In some cases, the observed positions of ligand atoms interacting with the proteins coincide with a region which is also frequently occupied by similar bonding partners in organic crystal structures, however, interaction geometries are also found which fall close to the limits of the ranges observed in the small molecule reference data. The information contained in the different composite crystal-field environments can be translated into rules which serve as guide-lines for automatic docking of small molecule fragments into the active site of proteins.

Binding Sites↗

Crystal structure of the N-terminal domain of the TyrR transcription factor responsible for gene regulation of aromatic amino acid biosynthesis and transport in Escherichia coli K12.

The X-ray structure of the N-terminal domain of TyrR has been solved to a resolution of 2.3 A. It reveals a modular protein containing an ACT domain, a connecting helix, a PAS domain and a C-terminal helix. Two dimers are present in the asymmetric unit with one monomer of each pair exhibiting a large rigid-body movement that results in a hinging around residue 74 of approximately 50 degrees . The structure of the dimer is discussed with reference to other transcription regulator proteins. Putative binding sites are identified for the aromatic amino acid cofactors.

Amino Acids, Aromatic↗

A gap-free, telomere-to-telomere chromosome-scale genome assembly of the mangrove red snapper, Lutjanus argentimaculatus.

The mangrove red snapper (Lutjanus argentimaculatus) is a commercially important marine fish species in the Indo-Pacific region. Despite its significant economic value for aquaculture, existing genomic resources remain fragmented, limiting the advancement of molecular breeding and functional genomic studies. Here, we present a gap-free, telomere-to-telomere (T2T) genome assembly of L. argentimaculatus, generated using a hybrid approach combining PacBio HiFi, Oxford Nanopore ultra-long reads and Hi-C technology. The resulting assembly comprises exactly 24 scaffolds spanning 1.03 Gb, perfectly matching the haploid chromosome number with a contig N50 of 46.17 Mb. Notably, this assembly resolves all physical gaps present in previous versions, achieving a BUSCO completeness score of 98.2%. Comprehensive genome annotation successfully predicted 23,167 protein-coding genes. Among these, 22,067 genes (95.25%) were functionally annotated across major public databases, including eggNOG, InterPro, and Swiss-Prot. Furthermore, structural analysis successfully identified 19 telomeres and 20 centromeres, validating the chromosomal integrity. This high-fidelity, gap-free reference genome provides a robust foundation for comparative genomics, population genetics, and the genetic improvement of Lutjanidae species.

Animals↗

mettannotator: a comprehensive and scalable Nextflow annotation pipeline for prokaryotic assemblies.

SUMMARY: In recent years, there has been a surge in prokaryotic genome assemblies, coming from both isolated organisms and environmental samples. These assemblies often include novel species that are poorly represented in reference databases creating a need for a tool that can annotate both well-described and novel taxa, and can run at scale. Here, we present mettannotator-a comprehensive, scalable Nextflow pipeline for prokaryotic genome annotation that identifies coding and noncoding regions, predicts protein functions, including antimicrobial resistance, and delineates gene clusters. The pipeline summarizes these results in a GFF (General Feature Format) file that can be easily utilized in downstream analysis or visualized using common genome browsers. Here, we show how it works on 200 genomes from 29 prokaryotic phyla, including isolate genomes and known and novel metagenome-assembled genomes, and present metrics on its performance in comparison to other tools. AVAILABILITY AND IMPLEMENTATION: The pipeline is written in Nextflow and Python and published under an open source Apache 2.0 licence. Instructions and source code can be accessed at https://github.com/EBI-Metagenomics/mettannotator. The pipeline is also available on WorkflowHub: https://workflowhub.eu/workflows/1069.

Software↗

Transterm: a database of mRNAs and translational control elements.

Transterm is a database that facilitates studies of translation and the translational control of protein synthesis. It contains a curated collection of elements in mRNAs that control translation, and biologically relevant mRNA regions extracted from GenBank. It is organised largely on a taxonomic basis with files and summaries for each species. Global patterns that may affect translation in particular species, for example bias in the context of initiation codons (Kozak's consensus or Shine-Dalgarno sequences) or termination codons, can be detected in the consensus and information content bias summaries. Several types of access are provided via a web browser interface. Transterm defined elements may be matched in a user's sequence or in the database. Alternatively, elements can be entered by the user to search specific sections of the database (for example, coding regions or 3' flanking regions or the 3'-UTRs) or the user's sequence. Each Transterm defined element has an associated biological description with references. The database is accessible at http://uther.otago.ac.nz/Transterm.html.

Animals↗