Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44Linked to original sources

Sequence analysis of the 144-kilobase accessory plasmid pSmeSM11a, isolated from a dominant Sinorhizobium meliloti strain identified during a long-term field release experiment.

The genome of Sinorhizobium meliloti type strain Rm1021 consists of three replicons: the chromosome and two megaplasmids, pSymA and pSymB. Additionally, many indigenous S. meliloti strains possess one or more smaller plasmids, which represent the accessory genome of this species. Here we describe the complete nucleotide sequence of an accessory plasmid, designated pSmeSM11a, that was isolated from a dominant indigenous S. meliloti subpopulation in the context of a long-term field release experiment with genetically modified S. meliloti strains. Sequence analysis of plasmid pSmeSM11a revealed that it is 144,170 bp long and has a mean G+C content of 59.5 mol%. Annotation of the sequence resulted in a total of 160 coding sequences. Functional predictions could be made for 43% of the genes, whereas 57% of the genes encode hypothetical or unknown gene products. Two plasmid replication modules, one belonging to the repABC replicon family and the other belonging to the plasmid type A replicator region family, were identified. Plasmid pSmeSM11a contains a mobilization (mob) module composed of the type IV secretion system-related genes traG and traA and a putative mobC gene. A large continuous region that is about 42 kb long is very similar to a corresponding region located on S. meliloti Rm1021 megaplasmid pSymA. Single-base-pair deletions in the homologous regions are responsible for frameshifts that result in nonparalogous coding sequences. Plasmid pSmeSM11a carries additional copies of the nodulation genes nodP and nodQ that are responsible for Nod factor sulfation. Furthermore, a tauD gene encoding a putative taurine dioxygenase was identified on pSmeSM11a. An acdS gene located on pSmeSM11a is the first example of such a gene in S. meliloti. The deduced acdS gene product is able to deaminate 1-aminocyclopropane-1-carboxylate and is proposed to be involved in reducing the phytohormone ethylene, thus influencing nodulation events. The presence of numerous insertion sequences suggests that these elements mediated acquisition of accessory plasmid modules.

Agrobacterium tumefaciens↗

Sequence and developmental expression of mRNA coding for a gap junction protein in Xenopus.

Cloned complementary DNAs representing the complete coding sequence for an embryonic gap junction protein in the frog Xenopus laevis have been isolated and sequenced. The cDNAs hybridize with an RNA of 1.5 kb that is first detected in gastrulating embryos and accumulates throughout gastrulation and neurulation. By the tailbud stage, the highest abundance of the transcript is found in the region containing ventroposterior endoderm and the rudiment of the liver. In the adult, transcripts are present in the lungs, alimentary tract organs, and kidneys, but are not detected in the brain, heart, body wall and skeletal muscles, spleen, or ovary. The gene encoding this embryonic gap junction protein is present in only one or a few copies in the frog genome. In vitro translation of RNA synthesized from the cDNA template produces a 30-kD protein, as predicted by the coding sequence. This product has extensive sequence similarity to mammalian gap junction proteins in its putative transmembrane and extracellular domains, but has diverged substantially in two of its intracellular domains.

Amino Acid Sequence↗

Heterochromatic sequences in a Drosophila whole-genome shotgun assembly.

BACKGROUND: Most eukaryotic genomes include a substantial repeat-rich fraction termed heterochromatin, which is concentrated in centric and telomeric regions. The repetitive nature of heterochromatic sequence makes it difficult to assemble and analyze. To better understand the heterochromatic component of the Drosophila melanogaster genome, we characterized and annotated portions of a whole-genome shotgun sequence assembly. RESULTS: WGS3, an improved whole-genome shotgun assembly, includes 20.7 Mb of draft-quality sequence not represented in the Release 3 sequence spanning the euchromatin. We annotated this sequence using the methods employed in the re-annotation of the Release 3 euchromatic sequence. This analysis predicted 297 protein-coding genes and six non-protein-coding genes, including known heterochromatic genes, and regions of similarity to known transposable elements. Bacterial artificial chromosome (BAC)-based fluorescence in situ hybridization analysis was used to correlate the genomic sequence with the cytogenetic map in order to refine the genomic definition of the centric heterochromatin; on the basis of our cytological definition, the annotated Release 3 euchromatic sequence extends into the centric heterochromatin on each chromosome arm. CONCLUSIONS: Whole-genome shotgun assembly produced a reliable draft-quality sequence of a significant part of the Drosophila heterochromatin. Annotation of this sequence defined the intron-exon structures of 30 known protein-coding genes and 267 protein-coding gene models. The cytogenetic mapping suggests that an additional 150 predicted genes are located in heterochromatin at the base of the Release 3 euchromatic sequence. Our analysis suggests strategies for improving the sequence and annotation of the heterochromatic portions of the Drosophila and other complex genomes.

Algorithms↗

cis-Regulatory inputs of the wnt8 gene in the sea urchin endomesoderm network.

Expression of the wnt8 gene is the key transcriptional motivator of an intercellular signaling loop which drives endomesoderm specification forward early in sea urchin embryogenesis. This gene was predicted by network perturbation analysis to be activated by inputs from the blimp1/krox gene, itself expressed zygotically in the endomesoderm during cleavage; and by a Tcf1/beta-catenin input. The implication is that zygotic expression of wnt8 is stimulated in neighboring cells by its own gene product, since reception of the Wnt8 ligand causes beta-catenin nuclearization. Here, the modular cis-regulatory system of the wnt8 gene of Strongylocentrotus purpuratus was characterized functionally, and shown to respond to blockade of both Blimp1/Krox and Tcf1/beta-catenin inputs just as does the endogenous gene. The genomic target sites for these factors were demonstrated by mutation in one of the cis-regulatory modules. The Tcf1/beta-catenin and Blimp1/Krox inputs are both necessary for normal endomesodermal expression mediated by this cis-regulatory module; thus, the genomic regulatory code underlying the predicted signaling loop thus resides in the wnt8 cis-regulatory sequence. In a second regulatory region, which initiates expression in micromere and macromere descendant cells early in cleavage, Tcf1 sites act to repress ectopic transcription in prospective ectoderm cells.

Animals↗

Sequence and analysis of bovine enteritic coronavirus (F15) genome. I. Sequence of the gene coding for the nucleocapsid protein; analysis of the predicted protein.

Sequences encoding the N protein of the bovine enteritic coronavirus-F15 strain (BECV-F15) have been cloned in PBR322 plasmid using cDNA produced by priming with oligo-dT on purified viral genomic RNA. Some 265 insert-containing clones were studied. Hybridization of these inserts with poly(A)+ RNA extracted from infected cells led to the conclusion that they were located at the 3'-end of the genome. After subcloning in M13 phage DNA, clones were sequenced by the Sanger technique. A 1,710-nucleotide sequence corresponding to the gene coding for the viral N-protein was established. It shows 2 overlapping open reading frames (ORF). The 3'-non-coding end of the gene has an 8-nucleotide sequence in common with the homologous genome areas of MHV, TGE and IBV viruses. This sequence may represent the polymerase RNA binding site. An upstream sequence surrounding the first AUG of the smaller ORF corresponds to a potentially functional initiation codon. The sequence of the primary translation product deduced from the DNA sequence predicts a polypeptide of 207 amino acids (22.9 Kd) with a high leucine (19.8%) content, possessing a hydrophobic N-terminal end. The larger ORF has a coding capacity of 448 amino acids (49.4 Kd), corresponding to the N-protein molecular weight. The deduced protein possesses 43 serine residues (9.6% of the total amino acid content) which may be phosphorylated and involved in N-protein/RNA binding. N-protein also has 5 regions with a high basic amino acid content. One of them is also serine-rich and has a strong homology site with MHV, TGE and IBV viruses. In the first part of the N-terminal, a 12-amino-acid sequence (PRWYFYYLGTGP) is highly conserved for BECV-F15, JHM, TGE and IBV viruses. BCV Mebus strain and BECV-F15 have only minor differences in their N-protein sequence.

Amino Acid Sequence↗

Comparisons of survival predictions using survival risk ratios based on International Classification of Diseases, Ninth Revision and Abbreviated Injury Scale trauma diagnosis codes.

BACKGROUND: We conducted a comparison of methods for predicting survival using survival risk ratios (SRRs), including new comparisons based on International Classification of Diseases, Ninth Revision (ICD-9) versus Abbreviated Injury Scale (AIS) six-digit codes. METHODS: From the Pennsylvania trauma center's registry, all direct trauma admissions were collected through June 22, 1999. Patients with no comorbid medical diagnoses and both ICD-9 and AIS injury codes were used for comparisons based on a single set of data. SRRs for ICD-9 and then for AIS diagnostic codes were each calculated two ways: from the survival rate of patients with each diagnosis and when each diagnosis was an isolated diagnosis. Probabilities of survival for the cohort were calculated using each set of SRRs by the multiplicative ICISS method and, where appropriate, the minimum SRR method. These prediction sets were then internally validated against actual survival by the Hosmer-Lemeshow goodness-of-fit statistic. RESULTS: The 41,364 patients had 1,224 different ICD-9 injury diagnoses in 32,261 combinations and 1,263 corresponding AIS injury diagnoses in 31,755 combinations, ranging from 1 to 27 injuries per patient. All conventional ICD-9-based combinations of SRRs and methods had better Hosmer-Lemeshow goodness-of-fit statistic fits than their AIS-based counterparts. The minimum SRR method produced better calibration than the multiplicative methods, presumably because it did not magnify inaccuracies in the SRRs that might occur with multiplication. CONCLUSION: Predictions of survival based on anatomic injury alone can be performed using ICD-9 codes, with no advantage from extra coding of AIS diagnoses. Predictions based on the single worst SRR were closer to actual outcomes than those based on multiplying SRRs.

Abbreviated Injury Scale↗

Organization of early region 1B of human adenovirus type 2: identification of four differentially spliced mRNAs.

The mRNAs from early region 1B of adenovirus type 2 have been studied by Northern blot, S1 nuclease, and cDNA analysis. Two novel mRNAs, designated 14S and 14.5S, have been observed in addition to the previously identified 9S, 13S, and 22S mRNAs. They are 1.26 and 1.31 kilobases long and differ from the 13S and 22S mRNAs in being composed of three exons instead of two. Their two terminal exons are the same as those present in the 13S mRNA, whereas the middle exon is unique to each of the two novel mRNA species. The structures of the 14S and 14.5S mRNAs allow the prediction of their coding capacities: both mRNA species, like the 22S and 13S mRNAs, contain an uninterrupted translational reading frame encoding a 21,000-molecular-weight (21K) polypeptide. The 14S mRNA can, in addition, encode a 16.5K polypeptide which shares N-terminal and C-terminal sequences with the 55K polypeptide, known to be encoded by the 22S mRNA. The 14.5S mRNA species encodes a hypothetical 9.2K polypeptide which has the same N terminus as the 55K polypeptide but a unique C terminus. The two mRNAs differ in their kinetics of appearance; the 14.5S mRNA is preferentially expressed late after infection in contrast to the 14S mRNA, which is present in approximately equal amounts early and late after infection. Taken together with previously published information the results suggest that early region 1B of adenovirus type 2 encodes five proteins in addition to virion polypeptide IX. These have predicted molecular weights of 55,000, 21,000, 16,500, 9,200, and 8,100.

Adenoviruses, Human↗

Curvature distribution in prokaryotic genomes.

DNA curvature is known to play a biological role in gene regulation, in particular, initiation of transcription. We applied the software CURVATURE based on the wedge model to predict whether promoter regions of certain prokaryotes may be characterized by higher intrinsic DNA curvature located within or upstream to these regions. The main purpose was to verify our earlier hypothesis that the DNA curvature plays a biological role in gene regulation in mesophilic as compared to hyperthermophilic prokaryotes, i.e., DNA curvature presumably has a functional adaptive significance determined by temperature selection. Therefore, we analyzed all available complete prokaryotic genomes. The analysis showed that there is a group of genomes with a relatively high average DNA curvature upstream of start of genes. Remarkably, all organisms of this group appeared to be mesophilic, which is a full confirmation of the former hypothesis. The conservative patterns of genomic curvature distribution across different mesophilic bacterial and archaeal genomes presented in this study provide a new, convincing indication that curved DNA is evolutionarily preserved and determined by temperature selection. Moreover, we found a rather peculiar property of hyperthermophilic prokaryotes: the coding regions are predicted to be significantly more curved than it would be expected from their dinucleotide composition.

DNA, Archaeal↗

RNA secondary structural alignment with conditional random fields.

MOTIVATION: The computational identification of non-coding RNA regions on the genome is currently receiving much attention. However, it is essentially harder than gene-finding problems for protein-coding regions because non-coding RNA sequences do not have strong statistical signals. Since comparative sequence analysis is effective for non-coding RNA detection, efficient computational methods are expected for structural alignment of RNA sequences. Several methods have been proposed to accomplish the structural alignment tasks for RNA sequences, and we found that one of the most important points is to estimate an accurate score matrix for calculating structural alignments. RESULTS: We propose a novel approach for RNA structural alignment based on conditional random fields (CRFs). Our approach has some specific features compared with previous methods in the sense that the parameters for structural alignment are estimated such that the model can most probably discriminate between correct alignments and incorrect alignments, and has the generalization ability so that a satisfiable score matrix can be obtained even with a small number of sample data without overfitting. Experimental results clearly show that the parameter estimation with CRFs can outperform all the other existing methods for structural alignments of RNA sequences. Furthermore, structural alignment search based on CRFs is more accurate for predicting non-coding RNA regions than the other scoring methods. These experimental results strongly support our discriminative method employing CRFs to estimate the score matrix parameters. AVAILABILITY: The program which is implemented in C++ is available at http://phmmts.dna.bio.keio.ac.jp/ under the GNU public license.

Base Sequence↗

Sequence analysis of the Ebola virus genome: organization, genetic elements, and comparison with the genome of Marburg virus.

Sequence analysis of the second through the sixth genes of the Ebola virus (EBO) genome indicates that it is organized similarly to rhabdoviruses and paramyxoviruses and is virtually the same as Marburg virus (MBG). In vitro translation experiments and predicted amino acid sequence comparisons showed that the order of the EBO genes is: 3'-NP-VP35-VP40-GP-VP30-VP24-L. The transcriptional start and stop (polyadenylation) signals are conserved and all contain the sequence 3'-UAAUU. Three base intergenic sequences are present between the NP and VP35 genes (3'-GAU) and VP40 and GP genes (3'-AGC), and a large intergenic sequence of 142 bases separates the VP30 and VP24 genes. Novel gene overlaps were found between the VP35 and VP40, the GP and VP30, and the VP24 and L genes. Overlaps are 20 or 18 bases in length and are limited to the conserved sequences determined for the transcriptional signals. Stem-and-loop structures were identified in the putative (+) leader RNA and at the 5' end of each mRNA. Hybridization studies showed that a small second mRNA is transcribed from the glycoprotein gene, and is produced by termination of transcription at an atypical polyadenylation signal located in the middle of the coding region. The predicted amino acid sequence of the glycoprotein contains an N-terminal signal peptide sequence, a hydrophobic anchor sequence, and 17 potential N-linked glycosylation sites. Alignment of predicted amino acid sequences showed that the structural proteins of EBO and MBG contain large regions of homology despite the absence of serologic cross-reactivity.

Amino Acid Sequence↗

Cloning and characterization of a Chlamydia psittaci gene coding for a protein localized in the inclusion membrane of infected cells.

Chlamydiae are obligate intracellular bacteria which occupy a non-acidified vacuole (the inclusion) throughout their developmental cycle. Little is known about events leading to the establishment and maintenance of the chlamydial inclusion membrane. To identify chlamydial proteins which are unique to the intracellular phase of the life cycle, an expression library of Chlamydia psittaci DNA was screened with convalescent antisera from infected animals and hyperimmune antisera generated against formalin-killed purified chlamydiae. Overlapping genomic clones were identified which expressed a 39 kDa protein only recognized by the convalescent sera. Sequence analysis of the clones identified two open reading frames (ORFs), one of which (ORF1) coded for a predicted 39 kDa gene product. The ORF1 sequence was amplified and fused to the malE gene of Escherichia coli and antisera were raised against the resulting fusion protein. Immunoblotting with these antisera demonstrated that the 39 kDa protein was present in lysates of infected cells and in reticulate bodies (RBs), but was at the limit of detection in lysates of purified C. psittaci elementary bodies. Fluorescence microscopy experiments demonstrated that this protein was localized in the inclusion membrane of infected HeLa cells, but was not detected on the developmental forms within the inclusion. Because the protein produced by ORF1 is deposited on the inclusion membrane of infected cells, this gene has been designated incA, (inclusion membrane protein A) and its gene product, IncA. In addition to the inclusion membrane, these antisera labelled structures that extended from the inclusion over the nucleus or into the cytoplasm of infected cells. Immunoblotting also demonstrated that IncA, in lysates of infected cells, had a migration pattern that seemed indicative of post-translational modification. This pattern was not observed in immunoblots of RBs or in the E. coli expressing IncA. Collectively, these data identify a chlamydial gene which codes for a protein that is released from RB and is localized in the inclusion membrane of infected cells.

Amino Acid Sequence↗

Structure and transforming function of transduced mutant alleles of the chicken c-myc gene.

A small retroviral vector carrying an oncogenic myc allele was isolated as a spontaneous variant (MH2E21) of avian oncovirus MH2. The MH2E21 genome, measuring only 2.3 kilobases, can be replicated like larger retroviral genomes and hence contains all cis-acting sequence elements essential for encapsidation and reverse transcription of retroviral RNA or for integration and transcription of proviral DNA. The MH2E21 genome contains 5' and 3' noncoding retroviral vector elements and a coding region comprising the first six codons of the viral gag gene and 417 v-myc codons. The gag-myc junction corresponds precisely to the presumed splice junction on subgenomic MH2 v-myc mRNA, the possible origin of MH2E21. Among the v-myc codons, the first 5 are derived from the noncoding 5' terminus of the second c-myc exon, and 412 codons correspond to the c-myc coding region. The predicted sequence of the MH2E21 protein product differs from that of the chicken c-myc protein by 11 additional amino-terminal residues and by 25 amino acid substitutions and a deletion of 4 residues within the shared domains. To investigate the functional significance of these structural changes, the MH2E21 genome was modified in vitro. The gag translational initiation codon was inactivated by oligonucleotide-directed mutagenesis. Furthermore, all but two of the missense mutations were reverted, and the deleted sequences were restored by replacing most of the MH2E21 v-myc allele by the corresponding segment of the CMII v-myc allele which is isogenic to c-myc in that region. The remaining two mutations have not been found in the v-myc alleles of avian oncoviruses MC29, CMII, and OK10. Like MH2 and MH2E21, modified MH2E21 (MH2E21m1c1) transforms avian embryo cells. Like c-myc, it encodes a 416-amino-acid protein initiated at the myc translational initiation codon. We conclude that neither major structural changes, such as in-frame fusion with virion genes or internal deletions, nor specific, if any, missense mutations of the c-myc coding region are necessary for activation of the basic oncogenic function of transduced myc alleles.

Alleles↗

Computational identification of Drosophila microRNA genes.

BACKGROUND: MicroRNAs (miRNAs) are a large family of 21-22 nucleotide non-coding RNAs with presumed post-transcriptional regulatory activity. Most miRNAs were identified by direct cloning of small RNAs, an approach that favors detection of abundant miRNAs. Three observations suggested that miRNA genes might be identified using a computational approach. First, miRNAs generally derive from precursor transcripts of 70-100 nucleotides with extended stem-loop structure. Second, miRNAs are usually highly conserved between the genomes of related species. Third, miRNAs display a characteristic pattern of evolutionary divergence. RESULTS: We developed an informatic procedure called 'miRseeker', which analyzed the completed euchromatic sequences of Drosophila melanogaster and D. pseudoobscura for conserved sequences that adopt an extended stem-loop structure and display a pattern of nucleotide divergence characteristic of known miRNAs. The sensitivity of this computational procedure was demonstrated by the presence of 75% (18/24) of previously identified Drosophila miRNAs within the top 124 candidates. In total, we identified 48 novel miRNA candidates that were strongly conserved in more distant insect, nematode, or vertebrate genomes. We verified expression for a total of 24 novel miRNA genes, including 20 of 27 candidates conserved in a third species and 4 of 11 high-scoring, Drosophila-specific candidates. Our analyses lead us to estimate that drosophilid genomes contain around 110 miRNA genes. CONCLUSIONS: Our computational strategy succeeded in identifying bona fide miRNA genes and suggests that miRNAs constitute nearly 1% of predicted protein-coding genes in Drosophila, a percentage similar to the percentage of miRNAs recently attributed to other metazoan genomes.

Animals↗

A 38 kb segment containing the cdc2 gene from the left arm of fission yeast chromosome II: sequence analysis and characterization of the genomic DNA and cDNAs encoded on the segment.

A genomic 38 kbp segment on the c1750 cosmid clone containing the cdc2 gene, located in the left arm of chromosome II from Schizosaccharomyces pombe, was sequenced. The segment was found to have five previously known genes, pht1, cdc2, his3, act1 and mei4. Among 11 coding sequences (CDSs) predicted by the gene finding software INTRON.PLOT., four CDSs, pi007, pi010, pi014 and pi016, had considerable similarity to 40S ribosomal protein, glycosyltransferase, cdc2-related protein kinase and alpha-1, 2-mannosyltransferase, respectively. Another unusually huge open reading frame (ORF) (pi011), consisting of 2233 amino acids, existed, having significant homology to alpha-amylase, granule-bound glycogen synthase and the Sz. pombe YS 1110 clone product at the N-terminal, middle and C-terminal regions, respectively. All the predicted 11 CDSs were experimentally analysed by RACE PCR. The sequencing of the RACE products revealed that there were two small overlaps at the 3' untranslated regions (UTRs) between pi004 and pi005 (17 bp) and between pi007 and pi008 (2 bp). The distances between 5' end of the 5'UTR and the putative translation initiation codon varied from 10 to 302 nucleotides (nt) among the nine CDSs successfully analysed by 5'-RACE. The expression level of each CDS on this clone was determined. Among the 16 genes on this clone, the previously determined genes, pht1, cdc2, his3 and act1, were found to be most highly expressed. Finally, cDNAs of all the newly identified genes were detected by RACE, proving the actual expression of these genes. The nucleotide sequence has been submitted to the EMBL database under Accession No. AB004534.

Base Sequence↗

Protein posttranslational modifications: the chemistry of proteome diversifications.

The diversity of distinct covalent forms of proteins (the proteome) greatly exceeds the number of proteins predicted by DNA coding capacities owing to directed posttranslational modifications. Enzymes dedicated to such protein modifications include 500 human protein kinases, 150 protein phosphatases, and 500 proteases. The major types of protein covalent modifications, such as phosphorylation, acetylation, glycosylation, methylation, and ubiquitylation, can be classified according to the type of amino acid side chain modified, the category of the modifying enzyme, and the extent of reversibility. Chemical events such as protein splicing, green fluorescent protein maturation, and proteasome autoactivations also represent posttranslational modifications. An understanding of the scope and pattern of the many posttranslational modifications in eukaryotic cells provides insight into the function and dynamics of proteome compositions.

Animals↗

Proteome analysis of Madrid E strain of Rickettsia prowazekii.

Rickettsia prowazekii, an obligate intracellular Gram-negative bacterium, is the etiologic agent of epidemic typhus. The threat of typhus as a biological weapon lies in its stability in the dried louse feces and in its infection by inhalation of an aerosol. Consequently, it is listed as a select agent and warrants more research to understand its pathogenesis. Although the genomic DNA sequence of strain Madrid E has been completed, the actual expression of the individual protein has not been investigated. In order to provide a global view of the expressed protein profile, the whole cell lysate of purified rickettsia (Madrid E strain) was reduced, alkylated, and digested with trypsin. The total digest was characterized by a two-dimensional liquid chromatography mass spectrometry system and analyzed with a modified version of the ProteomeX workstation. A total of 252 proteins out of 834 predicted protein-coding genes were identified, 238 proteins were identified by the detection of at least two unique peptides. Only 14 proteins were identified by the detection of one unique peptide in all three separate analyses. Among the 238 proteins identified by multiple unique peptides, 230 proteins were found in at least two of three separate analyses. The reproducible and convenient methodology and the information described here have provided a foundation for future proteome study of various R. prowazekii strains with different virulence.

Amino Acid Sequence↗

Quantitative structure-activity relationship study of bitter di- and tri-peptides including relationship with angiotensin I-converting enzyme inhibitory activity.

Bitterness represents a major challenge in industrial application of food protein hydrolysates or bioactive peptides and is a major factor that controls the flavor of formulated therapeutic products. The aim of this work was to apply quantitative structure-activity relationship modeling as a tool to determine the type and position of amino acids that contribute to bitterness of di- and tri-peptides. Datasets of bitter di- and tri-peptides were constructed using values from available literature, followed by modeling using partial least square (PLS) regression based on the three z-scores of 20 coded amino acids. Prediction models were validated using cross-validation and permutation tests. Results showed that a single-component model could explain 52 and 50% of the Y variance (bitterness threshold) of bitter di- and tri-peptides, respectively. Using PLS regression coefficients, it was determined that hydrophobic amino acids at the carboxyl-terminus and bulky amino acid residues adjacent to the carboxyl terminal are the major determinants of the intensity of bitterness of di- and tri-peptides. However, there was no significant (p > 0.05) correlation between bitterness of di- and tri-peptides and their angiotensin I-converting enzyme-inhibitory properties.

Amino Acid Sequence↗

KlSEC53 is an essential Kluyveromyces lactis gene and is homologous with the SEC53 gene of Saccharomyces cerevisiae.

Phosphomannomutase (PMM) is a key enzyme, which catalyses one of the first steps in the glycosylation pathway, the conversion of D-mannose-6-phosphate to D-mannose-1-phosphate. The latter is the substrate for the synthesis of GDP-mannose, which serves as the mannosyl donor for the glycosylation reactions in eukaryotic cells. In the yeast Saccharomyces cerevisiae PMM is encoded by the gene SEC53 (ScSEC53) and the deficiency of PMM activity leads to severe defects in both protein glycosylation and secretion. We report here on the isolation of the Kluyveromyces lactis SEC53 (KlSEC53) gene from a genomic library by virtue of its ability to complement a Saccharomyces cerevisiae sec53 mutation. The sequenced DNA fragment contained an open reading frame of 765 bp, coding for a predicted polypeptide, KlSec53p, of 254 amino acids. The KlSec53p displays a high degree of homology with phosphomannomutases from other yeast species, protozoans, plants and humans. Our results have demonstrated that KlSEC53 is the functional homologue of the ScSEC53 gene. Like ScSEC53, the KlSEC53 gene is essential for K. lactis cell viability. Phenotypic analysis of a K. lactis strain overexpressing the KlSEC53 gene revealed defects expected for impaired cell wall integrity.

Amino Acid Sequence↗