Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

mettannotator: a comprehensive and scalable Nextflow annotation pipeline for prokaryotic assemblies.

SUMMARY: In recent years, there has been a surge in prokaryotic genome assemblies, coming from both isolated organisms and environmental samples. These assemblies often include novel species that are poorly represented in reference databases creating a need for a tool that can annotate both well-described and novel taxa, and can run at scale. Here, we present mettannotator-a comprehensive, scalable Nextflow pipeline for prokaryotic genome annotation that identifies coding and noncoding regions, predicts protein functions, including antimicrobial resistance, and delineates gene clusters. The pipeline summarizes these results in a GFF (General Feature Format) file that can be easily utilized in downstream analysis or visualized using common genome browsers. Here, we show how it works on 200 genomes from 29 prokaryotic phyla, including isolate genomes and known and novel metagenome-assembled genomes, and present metrics on its performance in comparison to other tools. AVAILABILITY AND IMPLEMENTATION: The pipeline is written in Nextflow and Python and published under an open source Apache 2.0 licence. Instructions and source code can be accessed at https://github.com/EBI-Metagenomics/mettannotator. The pipeline is also available on WorkflowHub: https://workflowhub.eu/workflows/1069.

Software↗

Transterm: a database of mRNAs and translational control elements.

Transterm is a database that facilitates studies of translation and the translational control of protein synthesis. It contains a curated collection of elements in mRNAs that control translation, and biologically relevant mRNA regions extracted from GenBank. It is organised largely on a taxonomic basis with files and summaries for each species. Global patterns that may affect translation in particular species, for example bias in the context of initiation codons (Kozak's consensus or Shine-Dalgarno sequences) or termination codons, can be detected in the consensus and information content bias summaries. Several types of access are provided via a web browser interface. Transterm defined elements may be matched in a user's sequence or in the database. Alternatively, elements can be entered by the user to search specific sections of the database (for example, coding regions or 3' flanking regions or the 3'-UTRs) or the user's sequence. Each Transterm defined element has an associated biological description with references. The database is accessible at http://uther.otago.ac.nz/Transterm.html.

Animals↗

Signature-peptide approach to detecting proteins in complex mixtures.

The objective of the work presented in this paper was to test the concept that tryptic peptides may be used as analytical surrogates of the protein from which they were derived. Proteins in complex mixtures were digested with trypsin and classes of peptide fragments selected by affinity chromatography, lectin columns were used in this case. Affinity selected peptide mixtures were directly transferred to a high-resolution reversed-phase chromatography column and further resolved into fractions that were collected and subjected to matrix-assisted laser desorption ionization (MALDI) mass spectrometry. The presence of specific proteins was determined by identification of signature peptides in the mass spectra. Data are also presented that suggest proteins may be quantified as their signature peptides by using isotopically labeled internal standards. Isotope ratios of peptides were determined by MALDI mass spectrometry and used to determine the concentration of a peptide relative to that of the labeled internal standard. Peptides in tryptic digests were labeled by acetylation with acetyl N-hydroxysuccinimide while internal standard peptides were labeled with the trideuteroacetylated analogue. Advantages of this approach are that (i) it is easier to separate peptides than proteins, (ii) native structure of the protein does not have to be maintained during the analysis, (iii) structural variants do not interfere and (iv) putative proteins suggested from DNA databases can be recognized by using a signature peptide probe.

Amino Acid Sequence↗

[Analysis, identification and correction of some errors of model refseqs appeared in NCBI Human Gene Database by in silico cloning and experimental verification of novel human genes].

We found that human genome coding regions annotated by computers have different kinds of many errors in public domain through homologous BLAST of our cloned genes in non-redundant (nr) database, including insertions, deletions or mutations of one base pair or a segment in sequences at the cDNA level, or different permutation and combination of these errors. Basically, we use the three means for validating and identifying some errors of the model genes appeared in NCBI GENOME ANNOTATION PROJECT REFSEQS: (I) Evaluating the support degree of human EST clustering and draft human genome BLAST. (2) Preparation of chromosomal mapping of our verified genes and analysis of genomic organization of the genes. All of the exon/intron boundaries should be consistent with the GT/AG rule, and consensuses surrounding the splice boundaries should be found as well. (3) Experimental verification by RT-PCR of the in silico cloning genes and further by cDNA sequencing. And then we use the three means as reference: (1) Web searching or in silico cloning of the genes of different species, especially mouse and rat homologous genes, and thus judging the gene existence by ontology. (2) By using the released genes in public domain as standard, which should be highly homologous to our verified genes, especially the released human genes appeared in NCBI GENOME ANNOTATION PROJECT REFSEQS, we try to clone each a highly homologous complete gene similar to the released genes in public domain according to the strategy we developed in this paper. If we can not get it, our verified gene may be correct and the released gene in public domain may be wrong. (3) To find more evidence, we verified our cloned genes by RT-PCR or hybrid technique. Here we list some errors we found from NCBI GENOME ANNOTATION PROJECT REFSEQs: (1) Insert a base in the ORF by mistake which causes the frame shift of the coding amino acid. In detail, abase in the ORF of a gene is a redundant insertion, which causes a reading frame shift in the translation of an alternative protein, such as LOC124919 is wrong form of C17 orf32 (with mouse and rat orthologs determined by us). (2) Put together by mistake (with force). This is a wrong assembly of non-relating cDNA segment, such as LOC147007 is wrong form of C17orf32. (3) Mistakenly insert a base or one section of cDNA in the ORF which causes it ending beforehand, only coding cDNA sequence of N-terminal amino acids, incomplete. For example, LOC123722 is wrong form of SPRYD1, and even the human hypothetical gene LOC126250 or PDCD5 is wrong form of our PDCD5 (TFAR19). (4) Incomplete, only coding cDNA sequence of C-terminal amino acids. For example, human LOC149076 and mouse LOC230761 are wrong form of our verified human ZNF362 and mouse Zfp362, respectively. (5) Incomplete, only coding one section of coding protein cDNA sequence of correct gene ORF, lacking N-terminal and C-terminal amino acids sequence, and at the same time, mistakenly anticipates the first non-initiation codon amino acid of the incomplete protein amino acid as the initiation codon, e.g. anticipating L as M. For example, LOC200084 is wrong form of ZNF362. (6) Mistakenly insert a base or one section of cDNA in the ORF, wrongly causing unwanted termination codon before the insertion, so the coding protein lacks the first part of the amino acids. For example, the GenBank Acc. No. AL096883 ( LOCUS No. HS323M22B) is wrong form of an experimentally verified human NM_012263 with mouse ortholog of BC010510 determined. (7) It may regard the polluted genomic sequence as complete gene cDNA sequence and anticipate the so-called single exon gene, even the real one, only a small ORF in the very long single exon mRNA, while there really exists termination code in the same phase of the upper part of the ORF initiation code, no other characters accord with the gene's condition. For example, LOC91126 is wrong form of ZNF362. (8) The anticipated genes only have ORF which has no EST proofs on both terminal sides. Depending on this ORF, a complete gene cDNA with double support of EST and human genome (there are termination codes at the same phase of the upper part of ORF) which indicates the anticipated ORF reference sequence may be incorrect. For example, LOC164395 may be wrong form of novel human gene bankit4590055. (9) A similar but smaller protein-coding gene is anticipated in the range of the human genome sequence that has the support of EST experimental proof, so other new anticipated gene may be incorrect. For example, LOC167563 may be wrong form of CMYA5. However,these errors can be corrected or avoided by using our strategy. Here we give one example in detail: Comparision of the sequence SPRYD1 with human hypothetical gene LOC123722. The TAA bases in the position of 478-480 in LOC123722 cDNA is redundant, which causes a reading frame shift in the translation of an alternative protein. The redundancy of GTAAA of LOC123722 is not supported by our experimental clone,and is almost fully rejected by human EST alignment, and is shown as the next intron sequence by genomic GT/AG organization analysis. The verification of cDNA or genomic DNA sequence of SPRYD1 implies that LOC123722 has a wrong stop codon within its ORF because of the prediction program, thus being not complete cds. To sum up, by combining bioinformatics analyses with experimental verification, we have found that there are many errors of at least nine kinds appeared in NCBI GENOME ANNOTATION PROJECT REFSEQs through BLAST of our cloned genes in non-redundant database, and our strategy is helpful in correcting them, such as LOC14907, LOC200084 and LOC91126 (all of them should be ZNF362, but are three different kinds of wrong forms of ZNF362), three model reference sequences predicted from NCBI contig NT_004511 by automated computational analysis using gene prediction method, or such as LOC124919 and LOC147007 (both should be C17orf32, but are two different kinds of wrong forms of C17orf32), two model reference sequences predicted from NCBI contig NT_010808 by automated computational analysis using gene prediction method. Therefore, the correct identification and annotation of novel human genes may be still a heavy task, which can be finished within a long period of time. So human genome coding regions annotated by computer should be used with caution. The articles published in the past did not clearly point out the existence of mistakes in the NCBI human gene mode reference sequence. At the Seventh International Human Genome Conference held in April 2002, we first published the researching result on this aspect in the communication form of Posterly insert a base or one section of cDNA in the ORF, wrongly causing unwanted termination codon before the insertion, so the coding protein lacks the first part of the amino acids. For example, the GenBank Acc. No. AL096883 ( LOCUS No. HS323M22B) is wrong form of an experimentally verified human NM_012263 with mouse ortholog of BC010510 determined. (7) It may regard the polluted genomic sequence as complete gene cDNA sequence and anticipate the so-called single exon gene, even the real one, only a small ORF in the very long single exon mRNA, while there really exists termination code in the same phase of the upper part of the ORF initiation code, no other characters accord with the gene's condition. For example, LOC91126 is wrong form of ZNF362. (8) The anticipated genes only have ORF which has no EST proofs on both terminal sides. Depending on this ORF, a complete gene cDNA with double support of EST and human genome (there are termination codes at the same phase of the upper part of ORF) which indicates the anticipated ORF reference sequence may be incorrect. For example, LOC164395 may be wrong form of novel human gene bankit4590055. (9) A similar but smaller protein-coding gene is anticipated in the range of the human genome sequence that has the support of EST experimental proof, so other new anticipated gene may be incorrect. For example, LOC167563 may be wrong form of CMYA5. However, these errors can be corrected or avoided by using our strategy. Here we give one example in detail: Comparision of the sequence SPRYD1 with human hypothetical gene LOC123722. The TAA bases in the position of 478-480 in LOC123722 cDNA is redundant, which causes a reading frame shift in the translation of an alternative protein. The redundancy of GTAAA of LOC123722 is not supported by our experimental clone, and is almost fully rejected by human EST alignment, and is shown as the next intron sequence by genomic GT/AG organization analysis. The verification of cDNA or genomic DNA sequence of SPRYD1 implies that LOC123722 has a wrong stop codon within its ORF because of the prediction program, thus being not complete cds. To sum up, by combining bioinformatics analyses with experimental verification, we have found that there are many errors of at least nine kinds appeared in NCBI GENOME ANNOTATION PROJECT REFSEQs through BLAST of our cloned genes in non-redundant database, and our strategy is helpful in correcting them, such as LOC14907, LOC200084 and LOC91126 (all of them should be ZNF362, but are three different kinds of wrong forms of ZNF362), three model reference sequences predicted from NCBI contig NT_004511 by automated computational analysis using gene prediction method, or such as LOC124919 and LOC147007 (both should be C17orf32, but are two different kinds of wrong forms of C17orf32), two model reference sequences predicted from NCBI contig NT_010808 by automated computational analysis using gene prediction method. Therefore, the correct identification and annotation of novel human genes may be still a heavy task, which can be finished within a long period of time. So human genome coding regions annotated by computer should be used with caution. (ABSTRACT TRUNCATED)

Amino Acid Sequence↗

Proteomic analysis of the sarcosine-insoluble outer membrane fraction of the bacterial pathogen Bartonella henselae.

Bartonella henselae is an emerging zoonotic pathogen causing a wide range of disease manifestations in humans. In this study, we report on the analysis of the sarcosine-insoluble outer membrane fraction of B. henselae ATCC 49882 Houston-1 by one-dimensional sodium dodecyl sulfate-polyacrylamide gel electrophoresis (1-D SDS-PAGE) and two-dimensional nonequilibrium pH gradient polyacrylamide gel electrophoresis (2-D NEPHGE). Protein species were identified by matrix-assisted laser desorption/ionization-time of flight-mass spectrometry (MALDI-TOF-MS) and subsequent database query against the B. henselae genome sequence. Subcellular fractionation, application of the ionic detergent lauryl sarcosine, assessment of trypsin sensitivity, and heat modifiability of surface-exposed proteins represented valuable tools for the analysis of the outer membrane subproteome of B. henselae. 2-D NEPHGE was applied to display and catalogue a substantial number of proteins associated with the B. henselae sarcosine-insoluble outer membrane fraction, resulting in the establishment of a first 2-D reference map of this compartment. Thus, 53 distinct protein species associated with the outer membrane subproteome fraction were identified. This study provides novel insights into the membrane biology and the associated putative virulence factors of this pathogen of increasing medical importance.

Bacterial Proteins↗

Laboratory reference values for a group of captive Ball Pythons (Python regius)

OBJECTIVE: Laboratory reference values, including hematologic and serum biochemical variables, and oropharyngeal bacteria flora, were determined in a group of captive Ball Pythons (Python regius). ANIMALS: 20 adult Ball Pythons, weighing between 700 and 1,510 g, were allowed to acclimate at the recommended temperature range for the species (25 C night-time, up to 30 C daytime), then were evaluated for internal parasites and treated with appropriate medication prior to the start of the study. PROCEDURE: Hematologic values determined included WBC, hemoglobin, hematocrit, plasma protein, and differential cell count. Clinical biochemical analysis included determination of glucose, uric acid, calcium, phosphorus, total protein, alanine transaminase, alkaline phosphatase, and aspartate transaminase values. In addition to blood values, oropharyngeal swab specimens of the mouth were submitted for culture to determine the species of bacteria found in this population. Descriptive statistics were calculated for each hematologic and clinical biochemical value. Mean, SEM, and ranges were calculated. RESULTS: Hematologic values were similar to those reported in other snake species, except the hematocrit, which was lower. Clinical biochemical values different from those of other species were alkaline phosphatase activity, which was lower, and calcium and phosphorus concentrations, which were lower than values in other species. Bacteria isolated from the oropharynx were principally gram-negative organisms. CONCLUSION: Reference intervals reported in this study are important for establishing a database for comparative studies of Ball Pythons in other locations and under different husbandry conditions. CLINICAL RELEVANCE: Accumulated laboratory reference values will assist veterinarians in assessing the health status of Ball Pythons.

Alanine Transaminase↗

Evaluation of partial 16S ribosomal DNA sequencing for identification of nocardia species by using the MicroSeq 500 system with an expanded database.

Identification of clinically significant nocardiae to the species level is important in patient diagnosis and treatment. A study was performed to evaluate Nocardia species identification obtained by partial 16S ribosomal DNA (rDNA) sequencing by the MicroSeq 500 system with an expanded database. The expanded portion of the database was developed from partial 5' 16S rDNA sequences derived from 28 reference strains (from the American Type Culture Collection and the Japanese Collection of Microorganisms). The expanded MicroSeq 500 system was compared to (i). conventional identification obtained from a combination of growth characteristics with biochemical and drug susceptibility tests; (ii). molecular techniques involving restriction enzyme analysis (REA) of portions of the 16S rRNA and 65-kDa heat shock protein genes; and (iii). when necessary, sequencing of a 999-bp fragment of the 16S rRNA gene. An unknown isolate was identified as a particular species if the sequence obtained by partial 16S rDNA sequencing by the expanded MicroSeq 500 system was 99.0% similar to that of the reference strain. Ninety-four nocardiae representing 10 separate species were isolated from patient specimens and examined by using the three different methods. Sequencing of partial 16S rDNA by the expanded MicroSeq 500 system resulted in only 72% agreement with conventional methods for species identification and 90% agreement with the alternative molecular methods. Molecular methods for identification of Nocardia species provide more accurate and rapid results than the conventional methods using biochemical and susceptibility testing. With an expanded database, the MicroSeq 500 system for partial 16S rDNA was able to correctly identify the human pathogens N. brasiliensis, N. cyriacigeorgica, N. farcinica, N. nova, N. otitidiscaviarum, and N. veterana.

DNA, Ribosomal↗

Design and characterization of libraries of molecular fragments for use in NMR screening against protein targets.

We have designed four generations of a low molecular weight fragment library for use in NMR-based screening against protein targets. The library initially contained 723 fragments which were selected manually from the Available Chemicals Directory. A series of in silico filters and property calculations were developed to automate the selection process, allowing a larger database of 1.79 M available compounds to be searched for a further 357 compounds that were added to the library. A kinase binding pharmacophore was then derived to select 174 kinase-focused fragments. Finally, an additional 61 fragments were selected to increase the number of different pharmacophores represented within the library. All of the fragments added to the library passed quality checks to ensure they were suitable for the screening protocol, with appropriate solubility, purity, chemical stability, and unambiguous NMR spectrum. The successive generations of libraries have been characterized through analysis of structural properties (molecular weight, lipophilicity, polar surface area, number of rotatable bonds, and hydrogen-bonding potential) and by analyzing their pharmacophoric complexity. These calculations have been used to compare the fragment libraries with a drug-like reference set of compounds and a set of molecules that bind to protein active sites. In addition, an analysis of the overall results of screening the library against the ATP binding site of two protein targets (HSP90 and CDK2) reveals different patterns of fragment binding, demonstrating that the approach can find selective compounds that discriminate between related binding sites.

Algorithms↗

Human membrane transporter database: a Web-accessible relational database for drug transport studies and pharmacogenomics.

The human genome contains numerous genes that encode membrane transporters and related proteins. For drug discovery, development, and targeting, one needs to know which transporters play a role in drug disposition and effects. Moreover, genetic polymorphisms in human membrane transporters may contribute to interindividual differences in the response to drugs. Pharmacogenetics, and, on a genome-wide basis, pharmacogenomics, address the effect of genetic variants on an individual's response to drugs and xenobiotics. However, our knowledge of the relevant transporters is limited at present. To facilitate the study of drug transporters on a broad scale, including the use of microarray technology, we have constructed a human membrane transporter database (HMTD). Even though it is still largely incomplete, the database contains information on more than 250 human membrane transporters, such as sequence, gene family, structure, function, substrate, tissue distribution, and genetic disorders associated with transporter polymorphisms. Readers are invited to submit additional data. Implemented as a relational database, HMTD supports complex biological queries. Accessible through a Web browser user interface via Common Gateway Interface (CGI) and Java Database Connection (JDBC), HMTD also provides useful links and references, allowing interactive searching and downloading of data. Taking advantage of the features of an electronic journal, this paper serves as an interactive tutorial for using the database, which we expect to develop into a research tool.

Databases, Protein↗

Natural surfactant extract versus synthetic surfactant for neonatal respiratory distress syndrome.

BACKGROUND: This section is under preparation and will be included in the next issue. OBJECTIVES: To compare the effect of synthetic surfactant to natural surfactant in premature infants with established respiratory distress syndrome. SEARCH STRATEGY: Searches were made of the Oxford Database of Perinatal Trials, Medline (MeSH terms: pulmonary surfactant; limits: age groups, newborn infant; publication type, clinical trial), previous reviews including cross references, abstracts, conference and symposia proceedings, expert informants, and journal hand searching in the English language. SELECTION CRITERIA: Randomized controlled trials comparing administration of synthetic surfactants to administration of natural surfactant extracts in premature infants with respiratory distress syndrome were considered for this review. DATA COLLECTION AND ANALYSIS: Data regarding clinical outcomes including pneumothorax, patent ductus arteriosus, necrotizing enterocolitis, intraventricular hemorrhage (all intraventricular hemorrhage and severe intraventricular hemorrhage), chronic lung disease, retinopathy of prematurity, and mortality were excerpted by the primary reviewer (R. Soll). Data analysis was conducted according to the standards of the Neonatal Cochrane Review Group. MAIN RESULTS: The meta-analysis supports a significant reduction in the risk of pneumothorax (typical relative risk 0.68, 95% CI 0.56, 0.83; typical risk difference -0.04 95% CI -0.06, -0.02). No disadvantages to natural surfactant extract treatment are noted regarding other outcomes. A trend towards reduced mortality is noted in association with natural surfactant extract treatment. REVIEWER'S CONCLUSIONS: Both natural surfactant extracts and synthetic surfactant extracts are effective in the treatment of established respiratory distress syndrome. Comparative trials demonstrate greater early improvement in the requirement for ventilatory support and fewer pneumothoraces associated with natural surfactant extract treatment. On clinical grounds, natural surfactant extracts would seem to be the more desirable choice.

Biological Products↗

Flexible protein sequence patterns. A sensitive method to detect weak structural similarities.

The concept of a flexible protein sequence pattern is defined. In contrast to conventional pattern matching, template or sequence alignment methods, flexible patterns allow residue patterns typical of a complete protein fold to be developed in terms of residue positions (elements), separated by gaps of defined range. An efficient dynamic programming algorithm is presented to enable the best alignment(s) of a pattern with a sequence to be identified. The flexible pattern method is evaluated in detail by reference to the globin protein family, and by comparison to alignment techniques that exploit single sequence, multiple sequence and secondary structural information. A flexible pattern derived from seven globins aligned on structural criteria successfully discriminates all 345 globins from non-globins in the Protein Identification Resource database. Furthermore, a pattern that uses helical regions from just human alpha-haemoglobin identified 337 globins compared to 318 for the best non-pattern global alignment method. Patterns derived from successively fewer, yet more highly conserved positions in a structural alignment of seven globins show that as few as 38 residue positions (25 buried hydrophobic, 4 exposed and 9 others) may be used to uniquely identify the globin fold. The study suggests that flexible patterns gain discriminating power both by discarding regions known to vary within the protein family, and by defining gaps within specific ranges. Flexible patterns therefore provide a convenient and powerful bridge between regular expression pattern matching techniques and more conventional local and global sequence comparison algorithms.

Amino Acid Sequence↗

Comparative proteomic analysis of human whole saliva.

Human saliva performs a wide variety of biological functions that are critical for the maintenance of the oral health. Various functions include lubrication, buffering, antimicrobial protection, and the maintenance of mucosal integrity. In addition, whole saliva may be analysed for the diagnosis of human systemic diseases, since it can be readily collected and contains identifiable serum constituents. By using proteomic approach, we have established a reference proteome map of human whole saliva allowing for the resolution of greater than 200 protein spots in a single two-dimensional polyacrylamide gel. Fifty-four protein spots, comprised of 26 different proteins, were identifies using N-terminal sequencing, mass spectrometry, and/or computer matching with protein database. Ten proteins, whose levels were significantly different when bleeding had occurred in the oral cavity, were discussed in this study. These 10 proteins include alpha-1-antrypsin, apolipoprotein A-I, cystatin A, SA, SA-III, and SN, enolase I, hemoglobin beta-chain, thioredoxin peroxiredoxin B, as well as a prolactin-inducible protein. The proteomic approach identifies candidates from human whole saliva that may prove to be of diagnostic and therapeutic significance.

Adult↗

Identification of the alpha-aminoadipic semialdehyde synthase gene, which is defective in familial hyperlysinemia.

The first two steps in the mammalian lysine-degradation pathway are catalyzed by lysine-ketoglutarate reductase and saccharopine dehydrogenase, respectively, resulting in the conversion of lysine to alpha-aminoadipic semialdehyde. Defects in one or both of these activities result in familial hyperlysinemia, an autosomal recessive condition characterized by hyperlysinemia, lysinuria, and variable saccharopinuria. In yeast, lysine-ketoglutarate reductase and saccharopine dehydrogenase are encoded by the LYS1 and LYS9 genes, respectively, and we searched the available sequence databases for their human homologues. We identified a single cDNA that encoded an apparently bifunctional protein, with the N-terminal half similar to that of yeast LYS1 and with the C-terminal half similar to that of yeast LYS9. This bifunctional protein has previously been referred to as "alpha-aminoadipic semialdehyde synthase," and we have tentatively designated this gene "AASS." The AASS cDNA contains an open reading frame of 2,781 bp predicted to encode a 927-amino-acid-long protein. The gene has been sequenced and contains 24 exons scattered over 68 kb and maps to chromosome 7q31.3. Northern blot analysis revealed the presence of several transcripts in all tissues examined, with the highest expression occurring in the liver. We sequenced the genomic DNA from a single patient with hyperlysinemia (JJa). The patient is the product of a consanguineous mating and is homozygous for an out-of-frame 9-bp deletion in exon 15, which results in a premature stop codon at position 534 of the protein. On the basis of these and other results, we propose that AASS catalyzes the first two steps of the major lysine-degradation pathway in human cells and that inactivating mutations in the AASS gene are a cause of hyperlysinemia.

Amino Acid Sequence↗

Accuracy of urinalysis dipstick techniques in predicting significant proteinuria in pregnancy.

OBJECTIVE: To estimate the accuracy of point-of-care dipstick urinalysis in predicting significant proteinuria in pregnancy. DATA SOURCES: Literature from 1970 to February 2002 was identified via 1). general bibliographic databases, that is, MEDLINE and EMBASE, 2). Cochrane Library and relevant specialist register of the Cochrane Collaboration, and 3). checking the reference lists of known primary and review articles. METHODS OF STUDY SELECTION: Studies were selected if the accuracy of dipstick urinalysis techniques in predicting total protein excretion was estimated compared with a reference standard (laboratory estimation of protein excretion). The tests included visually read color-change dipsticks and automated dipstick urinalysis. Study selection, quality assessment, and data abstraction were performed independently and in duplicate. TABULATION, INTEGRATION, AND RESULTS: Data from selected studies were abstracted as 2 x 2 tables comparing the test result with the reference standard. Test accuracy was expressed as likelihood ratios. Summary likelihood ratios were generated as measures of diagnostic accuracy to determine posttest probabilities. The electronic search produced 1543 citations. After independent review of published articles, a total of 34 articles was obtained for further scrutiny, and 7 studies were considered eligible for inclusion in the review. The 6 studies evaluating visual dipstick urinalysis produced a pooled positive likelihood ratio of 3.48 (95% confidence interval 1.66, 7.27) and a pooled negative likelihood ratio of 0.6 (95% confidence interval 0.45, 0.8) for predicting 300 mg/24-hour proteinuria at the 1+ or greater threshold. CONCLUSION: The accuracy of dipstick urinalysis with a 1+ threshold in the prediction of significant proteinuria is poor and therefore of limited usefulness to the clinician. Accuracy may be improved at higher thresholds (greater than 1+ proteinuria), but available data are sparse and of poor methodological quality. Therefore, it is not possible to make meaningful inferences about accuracy at higher urine dipstick thresholds. There is an urgent need for research in this area of common obstetric practice.

Female↗

Computational analysis of evolution and conservation in a protein superfamily.

Many gene superfamilies have hundreds or thousands of members and hence pose a significant challenge when performing a large-scale phylogenetic analysis. Derivation of the most accurate alignment possible and inference of evolutionary relationships (with an appropriate measure of confidence) are significant "bottlenecks" in the process. A generally applicable strategy is outlined for identifying and aligning sequences, performing simple analysis of the resulting alignment, and inferring evolutionary relationships. Reference is made to the serpin superfamily. The 'partition cluster' method, a relatively rapid technique for extracting underlying associations from phylogenetic bootstrap trees, is also presented.

Computational Biology↗

Differentiation and characterization of enteroviruses by computer-assisted viral protein fingerprinting.

We have developed and standardized a computerized method for the typing and characterization of enteroviruses with radiolabeled viral protein fingerprints. Enteroviral proteins were radiolabeled with [35S]methionine during growth in cell culture and were then separated by polyacrylamide gel electrophoresis. The dried gel was scanned, and from the resulting computer image (which resembled an autoradiogram) protein patterns were computer extracted and stored in a database. The enterovirus database contained community and prototype strains belonging to 20 different enteroviral serotypes. Each serotype has a discrete protein pattern, and the most important pattern differences for determining each type are in the region of the viral capsid proteins VP1, VP2, and VP3. When the database was challenged with 148 clinical enterovirus strains, 144 (97%) were correctly identified by using the correlation coefficient as a quantitative measure of relatedness between two patterns. This method can identify a type in a single test and represents a practical alternative to virus neutralization because it is less expensive, is much faster (3 rather than 10 days), and does not rely on any virus-specific reagents. The results also show that most of the strains currently isolated from the community have protein patterns different from those of their older prototype strains. Viral protein fingerprinting is an evolving, dynamic system for the typing and characterization of enteroviruses. The method is appropriate for use in clinical virology and reference laboratories for the typing of enteroviruses, for the study of the epidemiology of enteroviruses, and for surveillance of enteroviruses.

Capsid↗

Computational analysis of composite regulatory elements.

Combinatorial regulation is a powerful mechanism for generating specificity in gene expression, and it is thought to play a pivotal role in the formation of the complex gene regulatory networks found in higher eukaryotes. The term "Composite Element" (CE) refers to a minimal functional unit where protein-DNA and protein-protein interactions contribute to a highly specific pattern of gene transcriptional regulation. Identification of composite elements will help to better understand gene regulation networks. Experimentally identified CEs are limited in number, and the currently available CE database COMPEL is based on such published information. Here, based on the statistical analysis of over-represented adjacent transcription factor binding sites, we describe a computational method to predict composite regulatory elements in genomic sequences. The algorithm proved to be efficient for extracting composite elements that had been experimentally confirmed and documented in the COMPEL database. Furthermore, putative new composite elements are predicted based on this method, and we have been able to confirm some of our predictions which are not included in the COMPEL database by searching published information.

Binding Sites↗

BioAfrica's HIV-1 proteomics resource: combining protein data with bioinformatics tools.

Most Internet online resources for investigating HIV biology contain either bioinformatics tools, protein information or sequence data. The objective of this study was to develop a comprehensive online proteomics resource that integrates bioinformatics with the latest information on HIV-1 protein structure, gene expression, post-transcriptional/post-translational modification, functional activity, and protein-macromolecule interactions. The BioAfrica HIV-1 Proteomics Resource http://bioafrica.mrc.ac.za/proteomics/index.html is a website that contains detailed information about the HIV-1 proteome and protease cleavage sites, as well as data-mining tools that can be used to manipulate and query protein sequence data, a BLAST tool for initiating structural analyses of HIV-1 proteins, and a proteomics tools directory. The Proteome section contains extensive data on each of 19 HIV-1 proteins, including their functional properties, a sample analysis of HIV-1HXB2, structural models and links to other online resources. The HIV-1 Protease Cleavage Sites section provides information on the position, subtype variation and genetic evolution of Gag, Gag-Pol and Nef cleavage sites. The HIV-1 Protein Data-mining Tool includes a set of 27 group M (subtypes A through K) reference sequences that can be used to assess the influence of genetic variation on immunological and functional domains of the protein. The BLAST Structure Tool identifies proteins with similar, experimentally determined topologies, and the Tools Directory provides a categorized list of websites and relevant software programs. This combined database and software repository is designed to facilitate the capture, retrieval and analysis of HIV-1 protein data, and to convert it into clinically useful information relating to the pathogenesis, transmission and therapeutic response of different HIV-1 variants. The HIV-1 Proteomics Resource is readily accessible through the BioAfrica website at: http://bioafrica.mrc.ac.za/proteomics/index.html.

Africa↗