Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

The identification and characterization of microsatellites in the compact genome of the Japanese pufferfish, Fugu rubripes: perspectives in functional and comparative genomic analyses.

Fugu rubripes (Fugu) has one of the smallest recorded vertebrate genomes and is an economic tool for comparative DNA sequence analysis. Initial characterization of 128 kb of Fugu DNA attributed the compactness of this genome, in part, to a sparseness of repetitive DNA sequence compared with mammalian genomic sequences. This paper describes a new and comprehensive analysis in which 501 theoretically possible microsatellites with a repeat unit of one to six bases were used to query two orders of magnitude more Fugu DNA (i.e. 11.338 Mb). A total of 6042 microsatellites were identified and categorized. In decreasing order, the 20 most frequently occurring microsatellites are AC, A, C, AGG, AG, AGC, AAT, AAAT, ACAG, ACGC, ATCC, AAC, ATC, AGGG, AAAG, AAG, AAAC, AT, CCG and TTAGGG. The 20 most frequently occurring microsatellites represent 81.79% of all microsatellites identified. Our results indicate that one microsatellite occurs every 1.876 kb of DNA in Fugu, 11.55% of the microsatellites are detected in open reading frames that are predicted protein coding regions. With respect to the proportion of microsatellites present in open reading frames and the total abundance (bp) of all microsatellites, the genome of Fugu is similar to the genome of many other vertebrate species. Previous estimates performed indicate that approximately 1% of many vertebrate genomes are comprized of microsatellite sequences. However, many differences prevail in the abundance and frequency of the individual microsatellite classes. Many of the frequently occurring microsatellites in Fugu are known to code in other species for regions in proteins such as transcription factors, whilst others are associated with known functions, such as transcription factor binding sites and form part of promoter regions in DNA sequences of genes. Therefore, it is likely that such repeats in genomes have a role in the evolution of genes, regulation of gene expression and consequently the evolution of species.

Animals↗

Identification, molecular cloning, and transcription analysis of the Choristoneura fumiferana nuclear polyhedrosis virus spindle-like protein gene.

The Choristoneura fumiferana nuclear polyhedrosis virus spindle-like protein (slp) gene has been identified and localized immediately downstream and in the same orientation as the CfMNPV DNA polymerase gene. The slp gene is 1101 bp long, predicted to code for a 366 amino acid (42.1 kDa) polypeptide. Transcriptional analysis revealed that the CfMNPV slp gene is expressed at late times postinfection, beginning at 24 hr postinfection and is most abundantly expressed after 36 hr. Transcription initiates within a single baculovirus consensus late start site sequence (GTAAG) at position -18 relative to the translation start codon. Based on amino acid comparisons, the CfMNPV gene is closely related to other similar baculovirus genes and distantly but recognizably related to the fusolin proteins of two entomopoxviruses. The conservation of amino acid sequence, glycosylation signals and specific domains throughout the protein suggest that this gene product may play an important role in insect DNA virus replication.

Amino Acid Sequence↗

Recent advances in molluscum contagiosum virus research.

Molluscum contagiosum virus (MCV) and variola virus (VAR) are the only two poxviruses that are specific for man. MCV causes skin tumors in humans and primarily in children and immunocompromised individuals. MCV is unable to replicate in tissue culture cells or animals. Recently, the DNA sequence of the 190 kbp MCV genome was reported by Senkevich et al. MCV was predicted to encode 163 proteins of which 103 were clearly related to those of smallpox virus. In contrast, it was found that MCV lacks 83 genes of VAR, including those involved in the suppression of the host response to infection, nucleotide biosynthesis, and cell proliferation. However, MCV possesses 59 genes predicted to code for novel proteins including MHC-class I, chemokine and glutathione peroxidase homologs not found in other poxviruses. The MCV genomic data allow the investigation of novel host defense mechanisms and provide new possibilities for the development of therapeutics for treatment and prevention of the MCV infection.

Animals↗

The POL1 gene from the fission yeast, Schizosaccharomyces pombe, shows conserved amino acid blocks specific for eukaryotic DNA polymerases alpha.

The POL1 gene of the fission yeast, Schizosaccharomyces pombe, was isolated using a POL1 gene probe from the budding yeast Saccharomyces cerevisiae, cloned and sequenced. This gene is unique and located on chromosome II. It includes a single 91 bp intron and is transcribed into a mRNA of about 4500 nucleotides. The predicted protein coded for by the S. pombe POL1 gene is 1405 amino acid long and its calculated molecular weight is about 160,000 daltons. This peptide contains seven amino acid blocks conserved among several DNA polymerases from different organisms and shares overall 37% and 34% identity with DNA polymerases alpha from S. cerevisiae and human cells, respectively. These results indicate that this gene codes for the S. pombe catalytic subunit of DNA polymerase alpha. The comparisons with human DNA polymerase alpha and with the budding yeast DNA polymerases alpha, delta and epsilon reveal conserved blocks of amino acids which are structurally and/or functionally specific only for eukaryotic alpha-type DNA polymerases.

Amino Acid Sequence↗

Molecular structure and genetic regulation of SFA, a gene responsible for resistance to formaldehyde in Saccharomyces cerevisiae, and characterization of its protein product.

A 3.7 kb DNA fragment of yeast chromosome IV has been sequenced that contains the SFA gene which, when present on a multi-copy plasmid in Saccharomyces cerevisiae, confers hyper-resistance to formaldehyde. The open reading frame of SFA is 1158 bp in size and encodes a polypeptide of 386 amino acids. The predicted protein shows strong homologies to several mammalian alcohol dehydrogenases and contains a sequence characteristic of binding sites for NAD. Overexpression of the SFA gene leads to enhanced consumption of formaldehyde, which is most probably the reason for the observed hyper-resistance phenotype. In sfa::LEU2 disruption mutants, sensitivity to formaldehyde is correlated with reduced degradation of the chemical. The SFA gene shares an 868 bp divergent promoter with UGX2 a gene of yet unknown function. Promoter deletion studies with a SFA promoter-lacZ gene fusion construct revealed negative interference on expression of SFA by upstream sequences. The upstream region between positions -145 and -172 is totally or partially responsible for control of inducibility of SFA by chemicals such as formaldehyde (FA), ethanol and methyl methanesulphonate. The 41 kDa SFA-encoded protein was purified from a hyper-resistant transformant; it oxidizes long-chain alcohols and, in the presence of glutathione, is able to oxidize FA. SFA is predicted to code for a long-chain alcohol dehydrogenase (glutathione-dependent formaldehyde dehydrogenase) of the yeast S. cerevisiae.

Alcohol Dehydrogenase↗

Genomic organization of a mouse MHC class II region including the H2-M and Lmp2 loci.

The region encompassing the Ma, Mb1, Mb2, and Lmp2 genes of the mouse class II major histocompatibility complex (MHC) was sequenced. Since this region contains clusters of genes required for efficient class I and class II antigen presentation, it was interesting to search for putative additional genes in the 21 kilobase gap between the Mb1 and Lmp2 genes. Computer predictions of coding regions and CpG islands, exon trapping experiments, and cross-species comparison with the corresponding human sequence indicate that no additional functional gene is present in that stretch. However, computer analysis revealed the possible existence of an alternative 3' exon for Mb1. Except for the fact that the mouse MHC contains two Mb genes, the genomic organization of the H2-M loci was found to be almost identical to the organization of the human HLA-DM genes. The promoter regions of the Ma and Mb genes also resemble classical class II promoters, containing typical S, X, and Y boxes. Like the human genes, the three H2-M genes displayed very limited polymorphism when we compared the cDNA sequences from six haplotypes. Finally, comparison of DMB with Mb1 and Mb2, both at the genomic level and in their coding regions, suggests that the Mb gene was recently duplicated, probably only in certain rodents.

Animals↗

Psychophysiological components of imagery.

McGuigan's neuromuscular model of information processing (1978a, 1978b, and 1989) was investigated by electrically recording eye movements (electro-oculograms), covert lip and preferred arm responses (electromyograms), and electroencephalograms. This model predicts that codes are generated as the lips are uniquely activated when processing words beginning with bilabial sounds like "p" or "b," as is the right arm to words like "pencil" that refer to its use. Twelve adult female participants selected for their high imagery ratings were asked to form images to three orally presented linguistic stimuli: the letter "p," the words "pencil" and "pasture," and to a control stimulus, the words "go blank." The following findings were significant beyond the 0.05 level: an increased covert lip response only to the letter "p," increased vertical eye activity to "p" and to the word "pencil," right arm response only to the word "pencil," and a decreased percentage of alpha waves from the right 02 lead only to the word "pasture." Since these covert responses uniquely occurred during specific imagery processes, it is inferred that they are components of neuromuscular circuits that function in accord with the model of information processing tested.

Adult↗

The bovine alpha-glucosidase gene: coding region, genomic structure, and mutations that cause bovine generalized glycogenosis.

We report here cDNA and genomic sequence of the bovine acidic alpha-glucosidase gene, from the initiation codon to the most 3' polyadenylation signal. The 2814-bp coding sequence predicts a 937-amino acid protein, which is highly conserved compared with the human alpha-glucosidase gene (86% and 83% identity respectively). The intron/exon boundaries are also conserved between the two species. Two mutations have been identified in Brahmans, and one in Shorthorns, that lead to generalized glycogenosis. All three mutations result in premature termination of translation. Evidence is also presented for a missense mutation segregating with the Brahman population, which is responsible for a 70-80% reduction in alpha-glucosidase activity.

Amino Acid Substitution↗

The 79,370-bp conjugative plasmid pB4 consists of an IncP-1beta backbone loaded with a chromate resistance transposon, the strA-strB streptomycin resistance gene pair, the oxacillinase gene bla(NPS-1), and a tripartite antibiotic efflux system of the resistance-nodulation-division family.

Plasmid pB4 is a conjugative antibiotic resistance plasmid, originally isolated from a microbial community growing in activated sludge, by means of an exogenous isolation method with Pseudomonas sp. B13 as recipient. We have determined the complete nucleotide sequence of pB4. The plasmid is 79,370 bp long and contains at least 81 complete coding regions. A suite of coding regions predicted to be involved in plasmid replication, plasmid maintenance, and conjugative transfer revealed significant similarity to the IncP-1beta backbone of R751. Four resistance gene regions comprising mobile genetic elements are inserted in the IncP-1beta backbone of pB4. The modular 'gene load' of pB4 includes (1) the novel transposon Tn 5719 containing genes characteristic of chromate resistance determinants, (2) the transposon Tn 5393c carrying the widespread streptomycin resistance gene pair strA-strB, (3) the beta-lactam antibiotic resistance gene bla(NPS-1) flanked by highly conserved sequences characteristic of integrons, and (4) a tripartite antibiotic resistance determinant comprising an efflux protein of the resistance-nodulation-division (RND) family, a periplasmic membrane fusion protein (MFP), and an outer membrane factor (OMF). The components of the RND-MFP-OMF efflux system showed the highest similarity to the products of the mexCD-oprJ determinant from the Pseudomonas aeruginosa chromosome. Functional analysis of the cloned resistance region from pB4 in Pseudomonas sp. B13 indicated that the RND-MFP-OMF efflux system conferred high-level resistance to erythromycin and roxithromycin resistance on the host strain. This is the first example of an RND-MFP-OMF-type antibiotic resistance determinant to be found in a plasmid genome. The global genetic organization of pB4 implies that its gene load might be disseminated between bacteria in different habitats by the combined action of the conjugation apparatus and the mobility of its component elements.

Bacterial Proteins↗

A gene for a Class II DNA photolyase from Oryza sativa: cloning of the cDNA by dilution-amplification.

Ultraviolet radiation induces the formation of two classes of photoproducts in DNA-the cyclobutane pyrimidine dimer (CPD) and the pyrimidine [6-4] pyrimidone photoproduct (6-4 product). Many organisms produce enzymes, termed photolyases, which specifically bind to these lesions and split them via a UV-A/blue light-dependent mechanism, thereby reversing the damage. These photolyases are specific for either CPDs or 6-4 products. Two classes of photolyases (class I and class II) repair CPDs. A gene that encodes a protein with class II CPD photolyase activity in vitro has been cloned from several plants including Arabidopsis thaliana, Cucumis sativus and Chlamydomonas reinhardtii. We report here the isolation of a homolog of this gene from rice (Oryza sativa), which was cloned on the basis of sequence similarity and PCR-based dilution-amplification. The cDNA comprises a very GC-rich (75%) 5; region, while the 3; portion has a GC content of 50%. This gene encodes a protein with CPD photolyase activity when expressed in E. coli. The CPD photolyase gene encodes at least two types of mRNA, formed by alternative splicing of exon 5. One of the mRNAs encodes an ORF for 506 amino acid residues, while the other is predicted to code for 364 amino acid residues. The two RNAs occur in about equal amounts in O. sativa cells.

Amino Acid Sequence↗

Gene content and organization of an 85-kb DNA segment from the genome of the phytopathogenic mollicute Spiroplasma kunkelii.

Spiroplasma kunkelii, the causative agent of corn stunt disease in maize (Zea maysL.), is a helical, cell wall-less prokaryote assigned to the class Mollicutes. As part of a project to sequence the entire S. kunkelii genome, we analyzed an 85-kb DNA segment from the pathogenic strain CR2-3x. This genome segment contains 101 ORFs and two tRNA genes. The majority of the ORFs code for predicted proteins that can be assigned to respective clusters of orthologous groups (COGs). These COGs cover diverse functional categories including genetic information storage and processing, cellular processes, and metabolism. The most notable gene cluster in this genome segment is a super-operon capable of encoding 24 ribosomal proteins. The organization of genes in this operon reflects the unique evolutionary position of the spiroplasma. Gene duplications, domain rearrangements, and frameshift mutations in the segment are interpreted as indicators of phase variation in the spiroplasma. To our knowledge, this is the first analysis of a large genome segment from a plant pathogenic spiroplasma.

Base Sequence↗

Mutational accessibility of essential genes on chromosome I(left) in Caenorhabditis elegans.

We have analyzed a region of approximately 5.4 million base pairs for mutations, which under standard laboratory conditions result in developmental arrest, sterility, or maternal-effect lethality in Caenorhabditis elegans. Lethal mutations were isolated, maintained, and genetically manipulated as homozygotes using sDp2--a duplication of the left half of chromosome I. All of the lethals and rearrangements used in this analysis were balanced by sDp2. Relatively low doses of mutagen, (approximately 15 mM ethylmethane sulfate; EMS), were used so as to limit the occurrence of second-site mutations, thus increasing the probability of recovering single nucleotide substitutions. Treatment of over 32,400 marked chromosomes resulted in 486 analyzed mutations. In this paper, we add 133 previously unidentified let genes, isolated in the EMS screens, and one let gene identified by a gamma-ray induced mutation, to our collection of 103 essential genes. We also recovered lethal alleles of genes for which visible mutants already existed. In total, eight deficiencies and alleles of 237 essential genes were identified. Eighty-nine of the previously unidentified let genes are represented by more than one lethal allele. Statistical analysis indicates a minimum estimate of 400 essential genes in the region of chromosome I balanced by sDp2. This region occupies approximately half of chromosome I, and contains over 1135 protein-coding genes predicted from the genomic sequence data. Thus, approximately one-third of the predicted genes are estimated to be essential. Of these approximately 60% are represented by lethal alleles. Less than 2% of the lethal-bearing strains recovered in our analysis, including the eight genetically definable deficiencies, carried more than one lethal mutation. Several screens were used to recover mutations for this analysis. Because all the mutations were isolated using the same balancer, under similar screening conditions, it was possible to compare intervals within the sDp2 region with each other. The fraction of essential genes that present relatively large targets for EMS was highest within the central cluster (dpy-5 to unc-13).

Animals↗

Molecular cloning and expression of rat hepatic neutral cholesteryl ester hydrolase.

The 1923 bp cDNA for rat hepatic cholesteryl ester hydrolase (CEH) was cloned by screening a lambda gt11 expression library with an oligonucleotide containing the consensus active site sequence for cholesteryl esterases. Expression of a fusion protein, cross-reacting with antibody to the purified liver CEH, was demonstrated by Western blot analysis. The cDNA was sequenced and found to have only 44% homology with pancreatic CEH. Although unique, the cDNA sequence exhibited much greater overall homology with liver carboxylesterases, in both coding and 5'/3' non-coding regions. In Northern blot analysis, the cDNA hybridized with a single band from liver mRNA but not with pancreatic mRNA. The 1.7 kb coding sequence, predicting a 62 kDa protein, was cloned into an Escherichia coli expression system with an inducible promoter and into COS-7 cells. Both expression systems produced a protein which comigrated with liver CEH (66 kDa) on SDS-PAGE and immunoreacted with antibodies to liver CEH on Western blots. Whereas the prokaryotic system produced an inactive protein, expression in COS-7 cells was accompanied by a 5-fold increase in CEH activity and a corresponding increase in immunoreactive protein.

Amino Acid Sequence↗

Molecular cloning and functional characterization of a GABA/betaine transporter from human kidney.

The human homologue of the canine GABA/betaine transporter (BGT-1) was isolated from a kidney inner medulla cDNA library. The coding sequence predicts a 614 amino acids protein with the typical features of neurotransmitter transporter family. The gene maps to chromosome 12p13 and, in addition to kidney, is also expressed in brain, liver, heart, skeletal muscle, and placenta. Functional studies reveal a Km = 20 microM for GABA transport and a coupling to Na+ and Cl- with a stoichiometry 3 Na+:2 Cl-:1 GABA. At 500 microM the GABA transport was inhibited by various compounds with the following potency order: quinidine > verapamil > phloretin > betaine.

Amino Acid Sequence↗

Codon usage in Plasmodium falciparum.

The codon frequencies used in 7874 codons from 17 sequences of Plasmodium falciparum have been examined. The frequency distribution is markedly biased. A and C occur with similar frequency in all positions but G is predominantly in the first base and T is predominantly in the last position. This information can be used to predict the coding strand and reading frame of P. falciparum genes.

Animals↗

The flp-1 propeptide is processed into multiple, highly similar FMRFamide-like peptides in Caenorhabditis elegans.

Previously, we described a gene, flp-1, that encodes seven FMRFamide-like peptides from two alternatively spliced transcripts in the nematode Caenorhabditis elegans. To determine whether all or a subset of the predicted peptides coded for by flp-1 are produced in vivo, we undertook the isolation of FMRFamide-like peptides from C. elegans. Six FLRFamide-containing peptides, all contained within the putative translation products of the flp-1 gene, were isolated from extracts of mixed stage animals. By quantitative PCR analysis of RNA from mixed stage animals, we found that the shorter transcript of flp-1 has a higher level of expression than the longer transcript.

Amino Acid Sequence↗

Evidence for two forms of murine beta-1,4-galactosyltransferase based on cloning studies.

We have isolated overlapping cDNA clones representing the full-length transcript (4038 base pairs) for murine beta-1,4-galactosyltransferase. The coding sequence predicts a membrane-bound glycoprotein with 3 distinct structural features: 1) a large, potentially glycosylated COOH-terminal domain (355 amino acids) which is positioned within the Golgi lumen and contains both the catalytic and alpha-lactalbumin binding site; 2) a single transmembrane domain (20 amino acids); and 3) a short NH2-terminal domain containing 2 Met residues, separated by 12 amino acids. The gene for murine beta-1,4-galactosyltransferase is unusual in that it specifies 2 mRNA transcripts which differ in length by about 200 base pairs. The longer transcript contains both Met residues found in the NH2-terminal domain; the shorter transcript contains only the downstream Met. These results predict that 2 related forms of beta-1,4-galactosyltransferase of 399 and 386 amino acids are synthesized as a consequence of alternative translation initiation. Both forms of the enzyme are identical in primary structure with the exception that the long form has an NH2-terminal extension of 13 amino acids which, in part, potentially encodes a cleavable signal sequence. The structural implications, topological distribution and potential biological significance of the 2 forms of the enzyme are discussed.

Amino Acid Sequence↗

Cloning and sequence analysis of cDNAs encoding human hippocampus N-methyl-D-aspartate receptor subunits: evidence for alternative RNA splicing.

Several cDNA clones encoding human N-methyl-D-aspartate receptor (hNR1) subunit polypeptides were isolated from a human hippocampus library. Degenerate oligodeoxyribonucleotide (oligo) primers based on the published rat NR1 (rNR1) amino acid (aa) sequence [K. Moriyoshi et al. Nature 354 (1991) 31-37] amplified a 0.7-kb fragment from a human hippocampus cDNA library, via the polymerase chain reaction (PCR). This fragment was used as a probe for subsequent hybridization screening. DNA sequence analysis of 28 plaque-purified clones indicated three distinct classes, designated hNR1-1, hNR1-2 and hNR1-3, presumably generated by alternative RNA splicing. One of these clones, hNR1-1(5A), was isolated as a full-length cDNA. The hNR1-2 and hNR1-3 cDNAs represented 66.8 and 98.9%, respectively, of the total aa coding information predicted for the polypeptides. The hNR1 cDNAs demonstrated an 84-90.8% nucleotide (nt) identity with the corresponding rodent cDNAs. The nt sequences of hNR1-1, hNR1-2 and hNR1-3 would encode 885-, 901- and 938-aa proteins, respectively, that have 99.1-99.8% identity with the corresponding rodent NR1 (roNR1) subunits. The changes between the predicted aa sequences of hNR1 and the corresponding roNR1 subunits are confined to the extracellular N-terminal regions. We have also identified two possible allelic variations of the hNR1-3 cDNA that result in aa substitutions in the extracellular N- and C-terminal regions. One of these naturally occurring aa variations is situated within a potential glutamate-binding site.

Alternative Splicing↗