Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

Human Plasma PeptideAtlas.

Peptide identifications of high probability from 28 LC-MS/MS human serum and plasma experiments from eight different laboratories, carried out in the context of the HUPO Plasma Proteome Project, were combined and mapped to the EnsEMBL human genome. The 6929 distinct observed peptides were mapped to approximately 960 different proteins. The resulting compendium of peptides and their associated samples, proteins, and genes is made publicly available as a reference for future research on human plasma.

Blood Proteins↗

Evaluation of protein fold comparison servers.

When a new protein structure has been determined, comparison with the database of known structures enables classification of its fold as new or belonging to a known class of proteins. This in turn may provide clues about the function of the protein. A large number of fold comparison programs have been developed, but they have never been subjected to a comprehensive and critical comparative analysis. Here we describe an evaluation of 11 publicly available, Web-based servers for automatic fold comparison. Both their functionality (e.g., user interface, presentation, and annotation of results) and their performance (i.e., how well established structural similarities are recognized) were assessed. The servers were subjected to a battery of performance tests covering a broad spectrum of folds as well as special cases, such as multidomain proteins, Calpha-only models, new folds, and NMR-based models. The CATH structural classification system was used as a reference. These tests revealed the strong and weak sides of each server. On the whole, CE, DALI, MATRAS, and VAST showed the best performance, but none of the servers achieved a 100% success rate. Where no structurally similar proteins are found by any individual server, it is recommended to try one or two other servers before any conclusions concerning the novelty of a fold are put on paper.

Computational Biology↗

Towards a reference map of Eimeria tenella sporozoite proteins by two-dimensional electrophoresis and mass spectrometry.

Eimeria tenella is a parasite of great importance as a disease causing agent in the poultry industry. Until recently, biological studies have focused on specific proteins, some of which play an important role in the parasite life cycle. Post-genomic studies will make it possible to understand the complexity of the parasites and their interactions with host cells. Here we present a systematic reference map of the proteins from E. tenella sporozoites. The proteins expressed at the sporozoite stage were resolved between isoelectric points 3-10 and 4-7. They were systematically identified using mass spectrometry and 16 known Eimeria sporozoite proteins were identified on two-dimensional maps. Peptide fragmentation data from mass spectrometry were compared to single and consensus expression sequence tags in databases and to the E. tenella genome (not annotated). Among the set of unknown proteins analysed, 12 new assignments were proposed on the basis of similarities with Apicomplexa proteins. In order to define sporozoite proteins as potential targets for coccidiosis therapy, proteins were studied according to their relative abundance and immunogenicity in the sporozoite. Immunoblots of sporozoite 2D maps with chicken sera were performed and approximately 50 proteins were defined as antigens. It was shown that abundance and immunogenicity are not related in the sporozoite stage. Perspectives of gene prediction and completion of the genome annotation by a proteomic approach is discussed.

Amino Acid Sequence↗

The molecular mechanism of cuproptosis and research progress in pancreatic diseases.

PURPOSE: Cuproptosis has been proven to be a novel mode of cell death, distinct from other types of cell death such as necrosis, ferroptosis, pyroptosis, and apoptosis. This study aims to systematically review the molecular mechanisms of cuproptosis in recent years and its research progress in pancreatic diseases. METHODS: By searching PubMed and Web of Science databases, 113 key literatures were included for thematic analysis, covering the molecular mechanism of cuproptosis and its role in the occurrence and development of pancreatic cancer, acute and chronic pancreatitis, diabetes, pancreatic cyst, pancreatic injury and pancreatic neuroendocrine tumor. RESULTS: Cuproptosis refers to the accumulation of copper ions in cells, which leads to instability of ferritin and aggregation of acylated proteins, resulting in oxidative stress-related cell death. Recent studies have shown that cuproptosis plays an important role in the occurrence and development of various pancreatic diseases, such as pancreatic cancer, acute and chronic pancreatitis, diabetes, pancreatic cysts, pancreatic injuries and pancreatic neuroendocrine tumor. The inducers of cuproptosis, such as disulfiram, chloroquinolones, and perilla phenols, alleviate pancreatic cancer by promoting cell cuproptosis. Copper chelators such as tetraethylenepentamine and tetrathiomolybdate promote the recovery of pancreatic injury by inhibiting cell cuproptosis. CONCLUSIONS: Cuproptosis plays a crucial role in the pathogenesis of pancreatic diseases. Further research on the cuproptosis pathway may become a potential target for the treatment of pancreatic diseases.

Animals↗

[Disease-causing mutations versus neutral polymorphism: use of bioinformatics and DNA diagnosis].

Molecular genetic diagnostics is available for increasing number of genetically determined diseases. A wide spectrum of mutations can be detected by laboratory methods. A mutation can be defined as a change in a specific DNA sequence when compared with the reference sequence published in the gene database. However, in some cases it is difficult to distinguish if the detected sequence variant is a causal mutation or a neutral (polymorphic) variation without any effect on phenotype. The interpretation of rare sequence variants of unknown significance detected in disease-causing genes becomes an increasingly important problem. Further analysis on DNA and on protein levels with the use of bioinformatics are needed to reveal the effect of rare sequence variants. Inherited complex disorders, for example rare hereditary forms of cancer diseases, represent a challenge to molecular geneticists. The identification of exact causal mutation directly responsible for the development of the disease and for the assessment of disease risk resulting from this genetic variation has further implications. Predictive genetic diagnostics allows identify relatives at high risk of genetically determined disease and use of targeted preventive and therapeutic approaches. In severe cases it allows also prenatal or pre-implantation diagnostics.

DNA Mutational Analysis↗

A multistep bioinformatic approach detects putative regulatory elements in gene promoters.

BACKGROUND: Searching for approximate patterns in large promoter sequences frequently produces an exceedingly high numbers of results. Our aim was to exploit biological knowledge for definition of a sheltered search space and of appropriate search parameters, in order to develop a method for identification of a tractable number of sequence motifs. RESULTS: Novel software (COOP) was developed for extraction of sequence motifs, based on clustering of exact or approximate patterns according to the frequency of their overlapping occurrences. Genomic sequences of 1 Kb upstream of 91 genes differentially expressed and/or encoding proteins with relevant function in adult human retina were analyzed. Methodology and results were tested by analysing 1,000 groups of putatively unrelated sequences, randomly selected among 17,156 human gene promoters. When applied to a sample of human promoters, the method identified 279 putative motifs frequently occurring in retina promoters sequences. Most of them are localized in the proximal portion of promoters, less variable in central region than in lateral regions and similar to known regulatory sequences. COOP software and reference manual are freely available upon request to the Authors. CONCLUSION: The approach described in this paper seems effective for identifying a tractable number of sequence motifs with putative regulatory role.

Algorithms↗

The archaea monophyly issue: A phylogeny of translational elongation factor G(2) sequences inferred from an optimized selection of alignment positions.

A global alignment of EF-G(2) sequences was corrected by reference to protein structure. The selection of characters eligible for construction of phylogenetic trees was optimized by searching for regions arising from the artifactual matching of sequence segments unique to different phylogenetic domains. The spurious matchings were identified by comparing all sections of the global alignment with a comprehensive inventory of significant binary alignments obtained by BLAST probing of the DNA and protein databases with representative EF-G(2) sequences. In three discrete alignment blocks (one in domain II and two in domain IV), the alignment of the bacterial sequences with those of Archaea-Eucarya was not retrieved by database probing with EF-G(2) sequences, and no EF-G homologue of the EF-2 sequence segments was detected by using partial EF-G(2) sequences as probes in BLAST/FASTA searches. The two domain IV regions (one of which comprises the ADP-ribosylatable site of EF-2) are almost certainly due to the artifactual alignment of insertion segments that are unique to Bacteria and to Archaea-Eucarya. Phylogenetic trees have been constructed from the global alignment after deselecting positions encompassing the unretrieved, spuriously aligned regions, as well as positions arising from misalignment of the G' and G" subdomain insertion segments flanking the "fifth" consensus motif of the G domain (AE varsson, 1995). The results show inconsistencies between trees inferred by alternative methods and alternative (DNA and protein) data sets with regard to Archaea being a monophyletic or paraphyletic grouping. Both maximum-likelihood and maximum-parsimony methods do not allow discrimination (by log-likelihood difference and difference in number of inferred substitutions) between the conflicting (monophyletic vs. paraphyletic Archaea) topologies. No specific EF-2 insertions (or terminal accretions) supporting a crenarchaeal-eucaryal clade are detectable in the new EF-G(2) sequence alignment.

Amino Acid Sequence↗

A two-dimensional proteome map of maize endosperm.

We have established a proteome reference map for maize (Zea mays L.) endosperm by means of two-dimensional gel electrophoresis and protein identification with LC-MS/MS analysis. This investigation focussed on proteins in major spots in a 4-7 pI range and 10-100 kDa M(r) range. Among the 632 protein spots processed, 496 were identified by matching against the NCBInr and ZMtuc-tus databases (using the SEQUEST software). Forty-two per cent of the proteins were identified against maize sequences, 23% against rice sequences and 21% against Arabidopsis sequences. Identified proteins were not only cytoplasmic but also nuclear, mitochondrial or amyloplastic. Metabolic processes, protein destination, protein synthesis, cell rescue, defense, cell death and ageing are the most abundant functional categories, comprising almost half of the 632 proteins analyzed in our study. This proteome map constitutes a powerful tool for physiological studies and is the first step for investigating the maize endosperm development.

Electrophoresis, Gel, Two-Dimensional↗

Separation and identification of soybean leaf proteins by two-dimensional gel electrophoresis and mass spectrometry.

To establish a proteomic reference map for soybean leaves, we separated and identified leaf proteins using two-dimensional polyacrylamide gel electrophoresis (2D-PAGE) and mass spectrometry (MS). Tryptic digests of 260 spots were subjected to peptide mass fingerprinting (PMF) by matrix-assisted laser desorption/ionization-time of flight (MALDI-TOF) MS. Fifty-three of these protein spots were identified by searching NCBInr and SwissProt databases using the Mascot search engine. Sixty-seven spots that were not identified by MALDI-TOF-MS analysis were analyzed with liquid chromatography tandem mass spectrometry (LC-MS/MS), and 66 of these spots were identified by searching against the NCBInr, SwissProt and expressed sequence tag (EST) databases. We have identified a total of 71 unique proteins. The majority of the identified leaf proteins are involved in energy metabolism. The results indicate that 2D-PAGE, combined with MALDI-TOF-MS and LC-MS/MS, is a sensitive and powerful technique for separation and identification of soybean leaf proteins. A summary of the identified proteins and their putative functions is discussed.

Amino Acid Sequence↗

Protein study of T and B acute lymphoblastic leukemia cell lines.

Two-dimensional polyacrylamide gel electrophoresis (2-D PAGE) was used to identify cellular proteins in T and B acute lymphoblastic leukemia (ALL) cell lines. Five lines, REH and BALL-1 of B-cell lineage, CCRF-CEM and HPB-ALL of T-cell lineage, and a normal Epstein-Barr virus (EBV)-transformed line of B-origin (SKLN1) were studied. The lines were immunophenotyped using flow cytometry and lineage associated monoclonal antibodies. Whole cell lysates of the cell lines were subjected to 2-D PAGE analyses. 2-D gels were analyzed with an image scanning computer and the qualitative as well as quantitative differences of the protein patterns were studied. Despite the great similarities in the patterns of the B- and T-gels, three proteins were unique to B-cell lines, while eight were unique to T-cell lines. Using cell lines is the first step toward identifying potential markers in ALL and can provide important information regarding the human ALL databases. Whether these proteins are definite markers for B- or T-ALL or are unique to the cell lines studied needs further exploration.

Algorithms↗

BTKbase: the mutation database for X-linked agammaglobulinemia.

X-linked agammaglobulinemia (XLA) is a hereditary immunodeficiency caused by mutations in the gene encoding Bruton tyrosine kinase (BTK). XLA patients have a decreased number of mature B cells and a lack of all immunoglobulin isotypes, resulting in susceptibility to severe bacterial infections. XLA-causing mutations are collected in a mutation database (BTKbase), which is available at http://bioinf.uta.fi/BTKbase. For each patient the following information is given (when available): the identification of the entry, a plain English description of the mutation followed by a reference, formal characterization of the mutation, and the various parameters from the patient. BTKbase is implemented with the MUTbase program suite, which provides an easy, interactive, and quality controlled submission of information to mutation databases. BTKbase version 8 lists mutation entries of 1,111 patients from 973 unrelated families showing 602 unique molecular events. The localization of the mutations on the gene and protein for BTK can be analyzed by clicking sequences on the web pages. The distribution of the mutations in the five structural domains is approximately proportional to the length of the domains, except for the Tec homology (TH) domain. The most frequently affected sites are CpG dinucleotides. The majority of the missense mutations are structural-disturbing Bruton tyrosine kinase (Btk) folding or decreasing stability. Many of the mutations affect functionally significant, conserved residues. The structural consequences of the mutations in all the domains have been studied based on crystallographic and nuclear magnetic resonance (NMR) structures as well as computer-aided molecular modeling.

Agammaglobulinemia↗

A sequence-based map of the nine genes of the human interleukin-1 cluster.

Six novel genes encoding proteins with the interleukin (IL)-1 fold have been identified recently. The classical family members are involved in inflammatory signaling. Previous work has placed the novel genes close to or within the same cluster as IL1A, IL1B, and IL1RN, which occupy an approximately 400-kb interval on chromosome 2. We have combined the incomplete public database sequence with our own sequence to generate a reference sequence and map that encompass all of the novel genes, allowing determination of the gene structures, precise localization of exons, and determination of distances between conventional SNP and microsatellite markers. Gene order from centromere to telomere is IL1A-IL1B-IL1F7-IL1F9-IL1F6-IL1F8-IL1F5-IL1F10-IL1RN, of which only IL1A, IL1B, and IL1F8 are transcribed towards the centromere. The gene order relates to the evolutionary relationship between the genes. Key features of exon boundaries are conserved. There is no evidence for other IL-1 family members within the cluster.

Amino Acid Sequence↗

Local polarity analysis: a sensitive method that discriminates between native proteins and incorrectly folded models.

The evaluation of calculated protein structures is an important step in the protein design cycle. Known criteria for this assessment of proteins are the polar and apolar, accessible and buried surface area, electrostatic interactions and other interactions between the protein atoms (e.g. H..O, S..S), atomic packing, analysis of amino acid environment and surface charge distribution. We show that a powerful test of accuracy of protein structure can be derived by analysing the water contact of atoms and additionally taking into account their polarity. On the basis of estimated reference values of the polar fraction of typical globular proteins with known structure (mean, SD and distribution), the evaluation of misfolded structures can be improved significantly. The reference values are derived by moving windows of different length (3-99 amino acid residues) over the amino acid sequence. Model proteins, which are included in the Brookhaven protein structure databank, deliberately misfolded proteins, hypothetical proteins and predicted protein structures are diagnosed as at least partially incorrectly folded. The local fault, mostly observed, is that polar groups are buried too frequently in the interior of the protein. The database-derived quantities are useful in screening the designed proteins prior to experimentation and may also be useful in the assessment of errors in the experimentally determined protein structures.

Chemical Phenomena↗

Virtual screening for anti-HIV-1 RT and anti-HIV-1 PR inhibitors from the Thai medicinal plants database: a combined docking with neural networks approach.

The virtual screening approach for docking small molecules into a known protein structure is a powerful tool for drug design. In this work, a combined docking and neural network approach, using a self-organizing map, has been developed and applied to screen anti-HIV-1 inhibitors for two targets, HIV-1 RT and HIV-1 PR, from active compounds available in the Thai Medicinal Plants Database. Based on nevirapine and calanolide A as reference structures in the HIV-1 RT binding site and XK-263 in the HIV-1 PR binding site, 2,684 compounds in the database were docked into the target enzymes. Self-organizing maps were then generated with respect to three types of pharmacophoric groups. The map of the reference structures were then superimposed on the feature maps of all screened compounds. Only the structures having similar features to the reference compounds were accepted. By using the SOMs, the number of candidates for HIV-1 RT was reduced to six and nine compounds consistent with nevirapine and calanolide A, respectively, as references. For the HIV-1 PR target, there are 135 screened compounds showed good agreement with the XK-263 feature map. These screened compounds will be further tested for their HIV-1 inhibitory affinities. The obtained results indicate that this combined method is clearly helpful to perform the successive screening and to reduce the analyzing step from AutoDock and scoring procedure.

Anti-HIV Agents↗

GeneQuiz: a workbench for sequence analysis.

We present the prototype of a software system, called GeneQuiz, for large-scale biological sequence analysis. The system was designed to meet the needs that arise in computational sequence analysis and our past experience with the analysis of 171 protein sequences of yeast chromosome III. We explain the cognitive challenges associated with this particular research activity and present our model of the sequence analysis process. The prototype system consists of two parts: (i) the database update and search system (driven by perl programs and rdb, a simple relational database engine also written in perl) and (ii) the visualization and browsing system (developed under C++/ET++). The principal design requirement for the first part was the complete automation of all repetitive actions: database updates, efficient sequence similarity searches and sampling of results in a uniform fashion. The user is then presented with "hit-lists" that summarize the results from heterogeneous database searches. The expert's primary task now simply becomes the further analysis of the candidate entries, where the problem is to extract adequate information about functional characteristics of the query protein rapidly. This second task is tremendously accelerated by a simple combination of the heterogeneous output into uniform relational tables and the provision of browsing mechanisms that give access to database records, sequence entries and alignment views. Indexing of molecular sequence databases provides fast retrieval of individual entries with the use of unique identifiers as well as browsing through databases using pre-existing cross-references. The presentation here covers an overview of the architecture of the system prototype and our experiences on its applicability in sequence analysis.(ABSTRACT TRUNCATED AT 250 WORDS)

Animals↗

Sensitivity of molecular docking to induced fit effects in influenza virus neuraminidase.

Many proteins undergo small side chain or even backbone movements on binding of different ligands into the same protein structure. This is known as induced fit and is potentially problematic for virtual screening of databases against protein targets. In this report we investigate the limits of the rigid protein approximation used by the docking program, GOLD, through cross-docking using protein structures of influenza neuraminidase. Neuraminidase is known to exhibit small but significant induced fit effects on ligand binding. Some neuraminidase crystal structures caused concern due to the bound ligand conformation and GOLD performed poorly on these complexes. A 'clean' set, which contained unique, unambiguous complexes, was defined. For this set, the lowest energy structure was correctly docked (i.e. RMSD < 1.5 A away from the crystal reference structure) in 84% of proteins, and the most promiscuous protein (1mwe) was able to dock all 15 ligands accurately including those that normally required an induced fit movement. This is considerably better than the 70% success rate seen with GOLD against general validation sets. Inclusion of specific water molecules involved in water-mediated hydrogen bonds did not significantly improve the docking performance for ligands that formed water-mediated contacts but it did prevent docking of ligands that displaced these waters. Our data supports the use of a single protein structure for virtual screening with GOLD in some applications involving induced fit effects, although care must be taken to identify the protein structure that performs best against a wide variety of ligands. The performance of GOLD was significantly better than the GOLD implementation of ChemScore and the reasons for this are discussed. Overall, GOLD has shown itself to be an extremely good, robust docking program for this system.

Algorithms↗

A new human hypervariable locus (K29) maps to the q37.3 region of chromosome 2 and reveals a fingerprint.

A human genomic library was screened with a 30-base oligomer corresponding to the 5' end of the human calretinin cDNA. A clone that contains a minisatellite composed of 21 imperfect repeats of a 37-bp sequence was isolated. The consensus (GAGGGAGGAACTGGGACGCGTGCATGTTTGCATTCTC) incidentally shares 14 consecutive matches with the oligomer used as a probe, and it was shown that the clone did not belong to the calretinin locus. The minisatellite, named K29, was used as a probe on Southern blots at high stringency. After HaeIII, MboI, or HinfI digestion, it detected a single hypervariable locus, with 65% heterozygosity among Caucasian individuals. The probe used at low stringency revealed a fingerprint, with an average of four bands in addition to the locus-specific pattern. Mendelian inheritance was assessed on pedigrees. The K29 minisatellite was mapped by in situ hybridization to the very end of the long arm of chromosome 2 (2q37.3 band), at close proximity of the Fra2J locus, and is referred to as the D2S88 locus in the genome database.

Base Sequence↗