Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 901 records · Page 50Linked to original sources

Genome sequence of the cyanobacterium Prochlorococcus marinus SS120, a nearly minimal oxyphototrophic genome.

Prochlorococcus marinus, the dominant photosynthetic organism in the ocean, is found in two main ecological forms: high-light-adapted genotypes in the upper part of the water column and low-light-adapted genotypes at the bottom of the illuminated layer. P. marinus SS120, the complete genome sequence reported here, is an extremely low-light-adapted form. The genome of P. marinus SS120 is composed of a single circular chromosome of 1,751,080 bp with an average G+C content of 36.4%. It contains 1,884 predicted protein-coding genes with an average size of 825 bp, a single rRNA operon, and 40 tRNA genes. Together with the 1.66-Mbp genome of P. marinus MED4, the genome of P. marinus SS120 is one of the two smallest genomes of a photosynthetic organism known to date. It lacks many genes that are involved in photosynthesis, DNA repair, solute uptake, intermediary metabolism, motility, phototaxis, and other functions that are conserved among other cyanobacteria. Systems of signal transduction and environmental stress response show a particularly drastic reduction in the number of components, even taking into account the small size of the SS120 genome. In contrast, housekeeping genes, which encode enzymes of amino acid, nucleotide, cofactor, and cell wall biosynthesis, are all present. Because of its remarkable compactness, the genome of P. marinus SS120 might approximate the minimal gene complement of a photosynthetic organism.

Adaptation, Physiological↗

cag, a pathogenicity island of Helicobacter pylori, encodes type I-specific and disease-associated virulence factors.

cagA, a gene that codes for an immunodominant antigen, is present only in Helicobacter pylori strains that are associated with severe forms of gastroduodenal disease (type I strains). We found that the genetic locus that contains cagA (cag) is part of a 40-kb DNA insertion that likely was acquired horizontally and integrated into the chromosomal glutamate racemase gene. This pathogenicity island is flanked by direct repeats of 31 bp. In some strains, cag is split into a right segment (cagI) and a left segment (cagII) by a novel insertion sequence (IS605). In a minority of H. pylori strains, cagI and cagII are separated by an intervening chromosomal sequence. Nucleotide sequencing of the 23,508 base pairs that form the cagI region and the extreme 3' end of the cagII region reveals the presence of 19 ORFs that code for proteins predicted to be mostly membrane associated with one gene (cagE), which is similar to the toxin-secretion gene of Bordetella pertussis, ptlC, and the transport systems required for plasmid transfer, including the virB4 gene of Agrobacterium tumefaciens. Transposon inactivation of several of the cagI genes abolishes induction of IL-8 expression in gastric epithelial cell lines. Thus, we believe the cag region may encode a novel H. pylori secretion system for the export of virulence determinants.

Antigens, Bacterial↗

The PR5K receptor protein kinase from Arabidopsis thaliana is structurally related to a family of plant defense proteins.

We have isolated an Arabidopsis thaliana gene that codes for a receptor related to antifungal pathogenesis-related (PR) proteins. The PR5K gene codes for a predicted 665-amino acid polypeptide that comprises an extracellular domain related to the PR5 proteins, a central transmembrane-spanning domain, and an intracellular protein-serine/threonine kinase. The extracellular domain of PR5K (PR5-like receptor kinase) is most highly related to acidic PR5 proteins that accumulate in the extracellular spaces of plants challenged with pathogenic microorganisms. The kinase domain of PR5K is related to a family of protein-serine/threonine kinases that are involved in the expression of self-incompatibility and disease resistance. PR5K transcripts accumulate at low levels in all tissues examined, although particularly high levels are present in roots and inflorescence stems. Treatments that induce authentic PR5 proteins had no effect on the level of PR5K transcripts, suggesting that the receptor forms part of a preexisting surveillance system. When the kinase domain of PR5K was expressed in Escherichia coli, the resulting polypeptide underwent autophosphorylation, consistent with its predicted enzyme activity. These results are consistent with PR5K encoding a functional receptor kinase. Moreover, the structural similarity between the extracellular domain of PR5K and the antimicrobial PR5- proteins suggests a possible interaction with common or related microbial targets.

Amino Acid Sequence↗

The biochemical and phenotypic characterization of Hho1p, the putative linker histone H1 of Saccharomyces cerevisiae.

There is currently no published report on the isolation and definitive identification of histone H1 in Saccharomyces cerevisiae. It was, however, recently shown that the yeast HHO1 gene codes for a predicted protein homologous to H1 of higher eukaryotes (Landsman, D. (1996) Trends Biochem. Sci. 21, 287-288; Ushinsky, S. C., Bussey, H. , Ahmed, A. A., Wang, Y., Friesen, J., Williams, B. A., and Storms, R. K. (1997) Yeast 13, 151-161), although there is no biochemical evidence that shows that Hho1p is, indeed, yeast histone H1. We showed that purified recombinant Hho1p (rHho1p) has electrophoretic and chromatographic properties similar to linker histones. The protein forms a stable ternary complex with a reconstituted core di-nucleosome in vitro at molar rHho1p:core ratios up to 1. Reconstitution of rHho1p with H1-stripped chromatin confers a kinetic pause at approximately 168 base pairs in the micrococcal nuclease digestion pattern of the chromatin. These results strongly suggest that Hho1p is a bona fide linker histone. We deleted the HHO1 gene and showed that the strain is viable and has no growth or mating defects. Hho1p is not required for telomeric silencing, basal transcriptional repression, or efficient sporulation. Unlike core histone mutations, a hho1Delta strain does not exhibit a Sin or Spt phenotype. The absence of Hho1p does not lead to a change in the nucleosome repeat length of bulk chromatin nor to differences in the in vivo micrococcal nuclease cleavage sites in individual genes as detected by primer extension mapping.

Amino Acid Sequence↗

Connexin43: a protein from rat heart homologous to a gap junction protein from liver.

Northern blot analysis of rat heart mRNA probed with a cDNA coding for the principal polypeptide of rat liver gap junctions demonstrated a 3.0-kb band. This band was observed only after hybridization and washing using low stringency conditions; high stringency conditions abolished the hybridization. A rat heart cDNA library was screened with the same cDNA probe under the permissive hybridization conditions, and a single positive clone identified and purified. The clone contained a 220-bp insert, which showed 55% homology to the original cDNA probe near the 5' end. The 220-bp cDNA was used to rescreen a heart cDNA library under high stringency conditions, and three additional cDNAs that together spanned 2,768 bp were isolated. This composite cDNA contained a single 1,146-bp open reading frame coding for a predicted polypeptide of 382 amino acids with a molecular mass of 43,036 D. Northern analysis of various rat tissues using this heart cDNA as probe showed hybridization to 3.0-kb bands in RNA isolated from heart, ovary, uterus, kidney, and lens epithelium. Comparisons of the predicted amino acid sequences for the two gap junction proteins isolated from heart and liver showed two regions of high homology (58 and 42%), and other regions of little or no homology. A model is presented which indicates that the conserved sequences correspond to transmembrane and extracellular regions of the junctional molecules, while the nonconserved sequences correspond to cytoplasmic regions. Since it has been shown previously that the original cDNA isolated from liver recognizes mRNAs in stomach, kidney, and brain, and it is shown here that the cDNA isolated from heart recognizes mRNAs in ovary, uterus, lens epithelium, and kidney, a nomenclature is proposed which avoids categorization by organ of origin. In this nomenclature, the homologous proteins in gap junctions would be called connexins, each distinguished by its predicted molecular mass in kilodaltons. The gap junction protein isolated from liver would then be called connexin32; from heart, connexin43.

Amino Acid Sequence↗

A unique point mutation in the PMP22 gene is associated with Charcot-Marie-Tooth disease and deafness.

Charcot-Marie-Tooth disease (CMT) with deafness is clinically distinct among the genetically heterogeneous group of CMT disorders. Molecular studies in a large family with autosomal dominant CMT and deafness have not been reported. The present molecular study involves a family with progressive features of CMT and deafness, originally reported by Kousseff et al. Genetic analysis of 70 individuals (31 affected, 28 unaffected, and 11 spouses) revealed linkage to markers on chromosome 17p11.2-p12, with a maximum LOD score of 9.01 for marker D17S1357 at a recombination fraction of .03. Haplotype analysis placed the CMT-deafness locus between markers D17S839 and D17S122, a approximately 0.6-Mb interval. This critical region lies within the CMT type 1A duplication region and excludes MYO15, a gene coding an unconventional myosin that causes a form of autosomal recessive deafness called DFNB3. Affected individuals from this family do not have the common 1.5-Mb duplication of CMT type 1A. Direct sequencing of the candidate peripheral myelin protein 22 (PMP22) gene detected a unique G-->C transversion in the heterozygous state in all affected individuals, at position 248 in coding exon 3, predicted to result in an Ala67Pro substitution in the second transmembrane domain of PMP22.

Amino Acid Sequence↗

A decision tree system for finding genes in DNA.

MORGAN is an integrated system for finding genes in vertebrate DNA sequences. MORGAN uses a variety of techniques to accomplish this task, the most distinctive of which is a decision tree classifier. The decision tree system is combined with new methods for identifying start codons, donor sites, and acceptor sites, and these are brought together in a frame-sensitive dynamic programming algorithm that finds the optimal segmentation of a DNA sequence into coding and noncoding regions (exons and introns). The optimal segmentation is dependent on a separate scoring function that takes a subsequence and assigns to it a score reflecting the probability that the sequence is an exon. The scoring functions in MORGAN are sets of decision trees that are combined to give a probability estimate. Experimental results on a database of 570 vertebrate DNA sequences show that MORGAN has excellent performance by many different measures. On a separate test set, it achieves an overall accuracy of 95 %, with a correlation coefficient of 0.78, and a sensitivity and specificity for coding bases of 83 % and 79%. In addition, MORGAN identifies 58% of coding exons exactly; i.e., both the beginning and end of the coding regions are predicted correctly. This paper describes the MORGAN system, including its decision tree routines and the algorithms for site recognition, and its performance on a benchmark database of vertebrate DNA.

Algorithms↗

Chicken interferon gene: cloning, expression, and analysis.

A gene encoding chicken interferon (ChIFN) was cloned from a cDNA library made from primary chick embryo cells that had been "aged" in vitro so as to produce copious amounts of IFN upon induction. The coding region is predicted to produce a signal peptide of 31 amino acids and a mature protein of 162 amino acids with a molecular weight of 18,957. There are four potential N-glycosylation sites and six cysteine residues. Three disulfide bonds are possible, with two being common to most mammalian type I IFNs. A motif of 10 amino acids surrounding Cys-137 is highly conserved: It shows 80% homology with mammalian type I IFNs, but only 30% with a reported fish IFN. The T-rich 3' UTR displays the canonical element AATAAA required for polyadenylation, and contains six repeats of the octamer CTATTTAT that may be involved in down-regulating translation. Northern blots demonstrate that the accumulation of ChIFN mRNA correlates with induction of ChIFN determined by bioassay. Biologically active protein was synthesized in transfected mouse L cells using mRNA prepared in vitro from the cloned sequence. This activity was neutralized by a monoclonal antibody prepared against purified ChIFN. The ChIFN gene shows sequence identity at the amino acid/nucleotide level with consensus mammalian IFNs as follows: alpha (24/23%), beta (20/24%), omega (23/43%), tau (20/43%), gamma (3/31%), and with flatfish IFN (16/35%). The conserved features of the predicted ChIFN protein and the general similarity of predicted secondary structure suggest a molecule that fits the five alpha-helix three-dimensional topology reported for type I mammalian IFNs.(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence↗

Cloning and expression of the chicken interferon-gamma gene.

We have cloned the gene for chicken interferon-gamma (ChIFN-gamma) from a cDNA expression library generated from a T cell line (CC8.1h) that produces high levels of IFN-gamma activity. CC8.1h constitutively produces IFN activity that shares physiochemical properties with mammalian IFN-gamma. ChIFN-gamma, when secreted by CC8.1h or expressed in transfected COS cells, is heat labile, inactivated by exposure to pH 2, and capable of inducing nitrite production by chicken macrophages. These properties clearly distinguish it from chicken and mammalian type I IFN. The ChIFN-gamma gene codes for a predicted mature protein of 145 amino acids with a molecular mass of 16.8 kD. There are two potential N-glycosylation sites located near the N terminus. ChIFN-gamma protein shares significant amino acid homology with mammalian IFN-gamma proteins; in particular it also contains the highly conserved motifs that are present in all mammalian IFN-gamma proteins. ChIFN-gamma is 35 and 32% identical to the equine and human counterparts, respectively, but shares only 15% homology with chicken type I IFN. These findings show that the emergence of the two principal types of IFN predates the divergence of avians and mammals that occurred some 350 million years ago.

Amino Acid Sequence↗

Probabilistic methods of identifying genes in prokaryotic genomes: connections to the HMM theory.

In this paper, we review developments in probabilistic methods of gene recognition in prokaryotic genomes with the emphasis on connections to the general theory of hidden Markov models (HMM). We show that the Bayesian method implemented in GeneMark, a frequently used gene-finding tool, can be augmented and reintroduced as a rigorous forward-backward (FB) algorithm for local posterior decoding described in the HMM theory. Another earlier developed method, prokaryotic GeneMark.hmm, uses a modification of the Viterbi algorithm for HMM with duration to identify the most likely global path through hidden functional states given the DNA sequence. GeneMark and GeneMark.hmm programs are worth using in concert for analysing prokaryotic DNA sequences that arguably do not follow any exact mathematical model. The new extension of GeneMark using the FB algorithm was implemented in the software program GeneMark.fba. Given the DNA sequence, this program determines an a posteriori probability for each nucleotide to belong to coding or non-coding region. Also, for any open reading frame (ORF), it assigns a score defined as a probabilistic measure of all paths through hidden states that traverse the ORF as a coding region. The prediction accuracy of GeneMark.fba determined in our tests was compared favourably to the accuracy of the initial (standard) GeneMark program. Comparison to the prokaryotic GeneMark.hmm has also demonstrated a certain, yet species-specific, degree of improvement in raw gene detection, ie detection of correct reading frame (and stop codon). The accuracy of exact gene prediction, which is concerned about precise prediction of gene start (which in a prokaryotic genome unambiguously defines the reading frame and stop codon, thus, the whole protein product), still remains more accurate in GeneMarkS, which uses more elaborate HMM to specifically address this task.

Algorithms↗

CONTRAfold: RNA secondary structure prediction without physics-based models.

MOTIVATION: For several decades, free energy minimization methods have been the dominant strategy for single sequence RNA secondary structure prediction. More recently, stochastic context-free grammars (SCFGs) have emerged as an alternative probabilistic methodology for modeling RNA structure. Unlike physics-based methods, which rely on thousands of experimentally-measured thermodynamic parameters, SCFGs use fully-automated statistical learning algorithms to derive model parameters. Despite this advantage, however, probabilistic methods have not replaced free energy minimization methods as the tool of choice for secondary structure prediction, as the accuracies of the best current SCFGs have yet to match those of the best physics-based models. RESULTS: In this paper, we present CONTRAfold, a novel secondary structure prediction method based on conditional log-linear models (CLLMs), a flexible class of probabilistic models which generalize upon SCFGs by using discriminative training and feature-rich scoring. In a series of cross-validation experiments, we show that grammar-based secondary structure prediction methods formulated as CLLMs consistently outperform their SCFG analogs. Furthermore, CONTRAfold, a CLLM incorporating most of the features found in typical thermodynamic models, achieves the highest single sequence prediction accuracies to date, outperforming currently available probabilistic and physics-based techniques. Our result thus closes the gap between probabilistic and thermodynamic models, demonstrating that statistical learning procedures provide an effective alternative to empirical measurement of thermodynamic parameters for RNA secondary structure prediction. AVAILABILITY: Source code for CONTRAfold is available at http://contra.stanford.edu/contrafold/.

Algorithms↗

Finding novel genes in bacterial communities isolated from the environment.

MOTIVATION: Novel sequencing techniques can give access to organisms that are difficult to cultivate using conventional methods. When applied to environmental samples, the data generated has some drawbacks, e.g. short length of assembled contigs, in-frame stop codons and frame shifts. Unfortunately, current gene finders cannot circumvent these difficulties. At the same time, the automated prediction of genes is a prerequisite for the increasing amount of genomic sequences to ensure progress in metagenomics. RESULTS: We introduce a novel gene finding algorithm that incorporates features overcoming the short length of the assembled contigs from environmental data, in-frame stop codons as well as frame shifts contained in bacterial sequences. The results show that by searching for sequence similarities in an environmental sample our algorithm is capable of detecting a high fraction of its gene content, depending on the species composition and the overall size of the sample. The method is valuable for hunting novel unknown genes that may be specific for the habitat where the sample is taken. Finally, we show that our algorithm can even exploit the limited information contained in the short reads generated by 454 technology for the prediction of protein coding genes. AVAILABILITY: The program is freely available upon request.

Algorithms↗

Neural bases of stereopsis across visual field of the alert macaque monkey.

Left and right retinal images of an object seen by the 2 eyes can occupy slightly disparate horizontal and/or vertical locations. The role of horizontal disparity (HD) in stereoscopic vision is well established, but the functional contribution of vertical disparity (VD) remains unclear. Various psychophysical studies have shown that HD and VD are used differently by the visual system depending on their location in the visual field, whether near the center of gaze or more peripheral. We show this horizontal/vertical distinction at the cellular level in monkey primary visual cortex (area V1). The range of VD encoding is reduced in central but not in the peripheral representation of the visual field. Moreover, neurons respond selectively to particular combinations of both types of disparities depending on the coded orientation as predicted by the disparity energy model. The preferred orientations of neurons near the fovea present a vertical bias that is well suited for stereopsis based on HD selectivity alone. In the periphery, instead, preferred orientations are radially biased, which allows a peripheral detector to convey the same depth signal based on either HD or VD. Such an organization has functional implications in both the perceptual and oculomotor domains.

Animals↗

Genome analysis of the glycosphingolipid-producing green alga tetraselmis sp. NKG400013.

Microalgae are gaining attention as sustainable resources for the production of valuable compounds, including biofuels, pigments, and bioactive metabolites. To support metabolic engineering and genome editing approaches aimed at enhancing these traits, high-quality genome assemblies are essential; however, genomic information remains limited for many microalgal lineages. Tetraselmis sp. NKG400013 is a green alga known for high glycosphingolipid accumulation with distinctive structural features. Here, we report a draft genome assembly of this strain generated using PacBio HiFi sequencing and transcriptome-supported annotation. The assembled genome spans 423.7 Mbp, with 74.5% repetitive sequences and 15,322 predicted protein-coding genes. Comparative analyses across 11 green algal species revealed a positive correlation between genome sizes and repeat contents, indicating that transposable element expansion, particularly long terminal repeat retrotransposons, has substantially contributed to genome enlargement in Tetraselmis. Genome-wide functional annotation and ortholog inference identified core enzymes required for glycosylceramide biosynthesis. Both sphingolipid Δ4 and Δ8 desaturases were identified in Tetraselmis and their coexistence suggests an expanded capacity for long-chain base modification that may underlie its distinctive glycosphingolipid profile. These results establish a genomic framework for understanding the high glycosphingolipid-producing capacity of NKG400013 and provide insights into the evolutionary diversification of sphingolipid metabolism in green algae.

Chlorophyta↗

Genetic characterization and cloning of mothers against dpp, a gene required for decapentaplegic function in Drosophila melanogaster.

The decapentaplegic (dpp) gene of Drosophila melanogaster encodes a growth factor that belongs to the transforming growth factor-beta (TGF-beta) superfamily and that plays a central role in multiple cell-cell signaling events throughout development. Through genetic screens we are seeking to identify other functions that act upstream, downstream or in concert with dpp to mediate its signaling role. We report here the genetic characterization and cloning of Mothers against dpp (Mad), a gene identified in two such screens. Mad loss-of-function mutations interact with dpp alleles to enhance embryonic dorsal-ventral patterning defects, as well as adult appendage defects, suggesting a role for Mad in mediating some aspect of dpp function. In support of this, homozygous Mad mutant animals exhibit defects in midgut morphogenesis, imaginal disk development and embryonic dorsal-ventral patterning that are very reminiscent of dpp mutant phenotypes. We cloned the Mad region and identified the Mad transcription unit through germline transformation rescue. We sequenced a Mad cDNA and identified three Mad point mutations that alter the coding information. The predicted MAD polypeptide lacks known protein motifs, but has strong sequence similarity to three polypeptides predicted from genomic sequence from the nematode Caenorhabditis elegans. Hence, MAD is a member of a novel, highly conserved protein family.

Amino Acid Sequence↗

Comparison of the tyrosine aminotransferase cDNA and genomic DNA sequences of normal mink and mink affected with tyrosinemia type II.

Type II tyrosinemia, designated Richner-Hanhart syndrome in humans, is a hereditary metabolic disorder with autosomal recessive inheritance characterized by a deficiency of tyrosine aminotransferase activity. Mutations occur in the human tyrosine aminotransferase gene, resulting in high levels of tyrosine and disease. Type II tyrosinemia occurs in mink, and our hypothesis was that it would also be associated with mutation(s) in the tyrosine aminotransferase gene. Therefore, the transcribed cDNA and the genomic tyrosine aminotransferase gene were sequenced from normal and affected mink. The gene extended over 11.9 kb and had 12 exons coding for a predicted 454-amino-acid protein with 93% homology with human tyrosine aminotransferase. FISH analysis mapped the gene to chromosome 8 using the Mandahl and Fredga (1975) nomenclature and chromosome 5 using the Christensen et al. (1996) nomenclature. The hypothesis was rejected because sequence analysis disclosed no mutations in either cDNA or introns that were associated with affected mink. This suggests that an unlinked gene regulatory mutation may be the cause of tyrosinemia in mink.

Amino Acid Sequence↗

Isolation, characterization and expression of the human Factor In the Germline alpha (FIGLA) gene in ovarian follicles and oocytes.

The Factor In the Germline alpha (FIGalpha) transcription factor regulates expression of the zona pellucida proteins ZP1, ZP2 and ZP3 and is essential for folliculogenesis in the mouse. Using the published mouse Figla sequence, BLAST searches identified a human chromosome 2 BAC clone with high sequence identity. Using PCR primers derived from this clone, amplicons derived from ovarian follicles and mature oocytes revealed 100% identity with the appropriate human BAC clone, the expected homology with the mouse Figla gene sequence, and homology on translation with the FIGalpha protein identified in the Japanese rice fish, medaka (Oryzias latipes). PCR expression profiling of this transcript revealed FIGLA mRNA expression in cDNA derived from ovarian follicles (5/5 samples from the primordial through to the secondary stage) mature oocytes (6/9 samples), and less frequently in preimplantation embryos (2/7 samples). Subsequent BLAST searches revealed the predicted full length coding sequence of the human FIGalpha protein which demonstrates 68 and 25% similarity overall to mouse and medaka proteins respectively, with 96 and 57% identity respectively within the basic helix-loop-helix region. This confirms our identification of the human homologue for this gene which maps to chromosome 2p12. Further work is required to understand its role in normal human oocyte development and the potential involvement in human infertility.

Alternative Splicing↗

Comparison of the amino acid sequence of the major immunogen from three serotypes of foot and mouth disease virus.

Cloned cDNA molecules from three serotypes of FMDV have been sequenced around the VP1-coding region. The predicted amino acid sequences for VP1 were compared with the published sequences and variable regions identified. The amino acid sequences were also analysed for hydrophilic regions. Two of the variable regions, numbered 129-160 and 193-204 overlapped hydrophilic regions, and were therefore identified as potentially immunogenic. These regions overlap regions shown by others to be immunogenic.

Amino Acid Sequence↗