Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 973 records · Page 54Linked to original sources

Genome sequences of six Enterobacter phages of the genus Karamvirus.

Here, we describe the genomes of six Enterobacter bacteriophages (phages) of the genus Karamvirus. The range of genome length, GC content, and number of predicted protein-coding sequences were, respectively, 171,900-175,675 bp, 39.54-39.82%, and 298-316. Genomic analysis indicates that these phages have a lytic lifestyle and are suitable for therapeutic use.

Enterobacter cloacae↗

Draft genomes of Limimaricola soesokkakensis TX01 isolated from the light organ of Anomalops katoptron in Indonesian waters.

The draft genome sequence of Limimaricola soesokkakensis TX01, isolated from the light organ of Anomalops katoptron, which lives in the region of the Banda Islands, Indonesia. The assembled genome is 4,041,441 base pairs in length, distributed across 69 contigs, with a GC content of 67.21%, and encodes 3,823 predicted protein-coding genes.

Limimaricola soesokkakensis↗

Wheat EST sequence assembly facilitates comparison of gene contents among plant species and discovery of novel genes.

Using a strategy requiring only modest computational resources, wheat expressed sequence tag (EST) sequences from various sources were assembled into contigs and compared with a nonredundant barley sequence assembly, with ESTs, with complete draft genome sequences of rice and Arabidopsis thaliana, and with ESTs from other plant species. These comparisons indicate that (i) wheat sequences available from public sources represent a substantial proportion of the diversity of wheat coding sequences, (ii) prediction of open reading frames in the whole genome sequence improves when supplemented with EST information from other species, (iii) a substantial number of candidates for novel genes that are unique to wheat or related species can be identified, and (iv) a smaller number of genes can be identified that are common to monocots and dicots but absent from Arabidopsis. The sequences in the last group may have been lost from Arabidopsis after descendance from a common ancestor. Examples of potential novel wheat genes and Triticeae-specific genes are presented.

Arabidopsis↗

The neuromuscular transform: the dynamic, nonlinear link between motor neuron firing patterns and muscle contraction in rhythmic behaviors.

The nervous system issues motor commands to muscles to generate behavior. All such commands must, however, pass through a filter that we call here the neuromuscular transform (NMT). The NMT transforms patterns of motor neuron firing to muscle contractions. This work is motivated by the fact that the NMT is far from being a straightforward, transparent link between motor neuron and muscle. The NMT is a dynamic, nonlinear, and modifiable filter. Consequently motor neuron firing translates to muscle contraction in a complex way. This complexity must be taken into account by the nervous system when issuing its motor commands, as well as by us when assessing their significance. This is the first of three papers in which we consider the properties and the functional role of the NMT. Physiologically, the motor neuron-muscle link comprises multiple steps of presynaptic and postsynaptic Ca(2+) elevation, transmitter release, and activation of the contractile machinery. The NMT formalizes all these into an overall input-output relation between patterns of motor neuron firing and shapes of muscle contractions. We develop here an analytic framework, essentially an elementary dynamical systems approach, with which we can study the global properties of the transformation. We analyze the principles that determine how different firing patterns are transformed to contractions, and different parameters of the former to parameters of the latter. The key properties of the NMT are its nonlinearity and its time dependence, relative to the time scale of the firing pattern. We then discuss issues of neuromuscular prediction, control, and coding. Does the firing pattern contain a code by means of which particular parameters of motor neuron firing control particular parameters of muscle contraction? What information must the motor neuron, and the nervous system generally, have about the periphery to be able to control it effectively? We focus here particularly on cyclical, rhythmic contractions which reveal the principles particularly clearly. Where possible, we illustrate the principles in an experimentally advantageous model system, the accessory radula closer (ARC)-opener neuromuscular system of Aplysia. In the following papers, we use the framework developed here to examine how the properties of the NMT govern functional performance in different rhythmic behaviors that the nervous system may command.

Animals↗

Somatic mutation and antibody diversity.

Nucleotide insertions or deletions determine novel amino acid sequences at the VH-D and D-JH junction sites. Since these cannot be predicted by known coding genes, they are regarded as a form of somatic mutagenesis. A second type of somatic mutation in Ig structural genes are the stochastic base substitutions that have now been found in both V region coding sequences and noncoding flanking sequences. It has been proposed that this form of mutagenesis may be related to the gene rearrangement process as well. A third type of mutagenesis has been associated with Ig CH switching. In this case, a higher frequency of substitutions has been observed in VH regions associated with C gamma and C alpha. A mechanism for the origin of this phenomenon is not known. An unexpectedly high rate of mutagenesis affecting Ig V genes has been observed to occur spontaneously in vitro.

Amino Acid Sequence↗

Transcriptome analysis of antigenic variation in Plasmodium falciparum--var silencing is not dependent on antisense RNA.

BACKGROUND: Plasmodium falciparum, the causative agent of the most severe form of malaria, undergoes antigenic variation through successive presentation of a family of antigens on the surface of parasitized erythrocytes. These antigens, known as Plasmodium falciparum erythrocyte membrane protein 1 (PfEMP1) proteins, are subject to a mutually exclusive expression system, and are encoded by the multigene var family. The mechanism whereby inactive var genes are silenced is poorly understood. To investigate transcriptional features of this mechanism, we conducted a microarray analysis of parasites that were selected to express different var genes by adhesion to chondroitin sulfate A (CSA) or CD36. RESULTS: In addition to oligonucleotides for all predicted protein-coding genes, oligonucleotide probes specific to each known var gene of the FCR3 background were designed and added to the microarray, as well as tiled sense and antisense probes for a subset of var genes. In parasites selected for adhesion to CSA, one full-length var gene (var2csa) was strongly upregulated, as were sense RNA molecules emanating from the 3' end of a limited subset of other var genes. No global relationship between sense and antisense production of var genes was observed, but notably, some var genes had coincident high levels of both antisense and sense transcript. CONCLUSION: Mutually exclusive expression of PfEMP1 proteins results from transcriptional silencing of non-expressed var genes. The distribution of steady-state sense and antisense RNA at var loci are not consistent with a silencing mechanism based on antisense silencing of inactive var genes. Silencing of var loci is also associated with altered regulation of genes distal to var loci.

Animals↗

Long-read transcriptomics corrects Trichomonas vaginalis intron annotations and refines transcript-end features.

BACKGROUND: Trichomonas vaginalis causes the most prevalent non-viral sexually transmitted infection worldwide. Despite its large genome (181.5 Mb; 36,310 predicted protein-coding genes in NYU_TvagG3_2), intron annotations remain limited and inconsistently validated. A recent short-read RNA-seq study reported 63 putative active introns, but short reads can misassign splice boundaries and cannot resolve complete transcript structures. METHODS: We integrated Oxford Nanopore direct RNA sequencing (DRS), ONT cDNA long-read sequencing, and Illumina RNA-seq to refine intron annotations, transcript-end features, and UTR boundaries in T. vaginalis. Candidate introns were validated by targeted PCR and Sanger sequencing, and representative splicing events were further assessed using public SRA datasets. RESULTS: Starting from 31 historically annotated introns, motif-guided long-read screening and orthogonal validation identified 17 additional validated introns, increasing the curated set to 48 confirmed introns. Among these 17 events, three were previously unrecognized in the current NYU_TvagG3_2 reference annotation. We also corrected five reported loci, including two false-positive introns, two splice-coordinate misannotations, and one gene-sequence error. DRS further supported transcript termination site mapping, UAAA polyadenylation-signal profiling relative to poly(A) addition sites, and single-molecule poly(A)-tail estimation. StringTie mixed-mode assemblies provided updated UTR boundaries for intron-bearing transcripts and transcripts without curated introns. CONCLUSIONS: This study provides a rigorously validated, long-read-refined resource of intron annotations, UTR boundaries, and UAAA-guided transcript-end features for T. vaginalis, together with a reproducible workflow for non-model protists. These refinements improve the current reference annotation and support future studies of functional genomics, parasite biology, pathogenesis, and diagnostic development.

Trichomonas vaginalis↗

Band 6 protein, a major constituent of desmosomes from stratified epithelia, is a novel member of the armadillo multigene family.

Desmosomes are intercellular adhering junctions characteristic of epithelial cells. Several constitutive proteins--desmoplakin, plakoglobin and the transmembrane glycoproteins desmoglein and desmocollin--have been identified as fundamental constituents of desmosomes in all tissues. A number of additional and cell type-specific constituents also contribute to desmosomal plaque formation. Among these proteins is the band 6 polypeptide (B6P). This positively charged, non-glycosylated protein is a major constituent of the plaque in stratified and complex glandular epithelia. Using an overlay assay we show that purified keratins bind in vitro to B6P. Thus B6P may play a role in ordering intermediate filament networks of adjacent epithelial cells. To characterize the structure of B6P in the desmosome we have isolated cDNA clones representing the entire coding sequence. The predicted amino acid sequence of human B6P shows strong sequence homology with a murine p120 protein, which is a substrate of protein tyrosine kinase receptors and of p60v-src. P120 and B6P show amino-terminal domains differing distinctly in length and sequence. These are followed in both proteins by 460 residues that display a series of imperfect repeats corresponding to the repeats in the cadherin binding proteins armadillo, plakoglobin and beta-catenin. Over this repeat region B6P and p120 share 33% sequence identity (54% similarity). These sequence characteristics define B6P as a novel member of the armadillo multigene family and raise the question of whether the structural proteins B6P, plakoglobin, beta-catenin and armadillo share some function. Since armadillo, plakoglobin, beta-catenin and p120 seem involved in signal transduction this may also hold for B6P. The amino-terminal region of B6P (residues 1 to 263) shows no significant homology to any known protein sequence. It may therefore be involved in unique functions of B6P.

Amino Acid Sequence↗

Sequence and expression of the CAPA/CAP2b gene in the tobacco hawkmoth, Manduca sexta.

The gene coding for cardioacceleratory peptide 2b (CAP2b; pELYAFPRV) has been isolated and sequenced from the moth Manduca sexta (GenBank accession #AY649544). Because of its significant homology to the CAPA gene in Drosophila melanogaster, this gene is called the Manduca CAPA gene. The Manduca CAPA gene is 958 nucleotides long with 29 untranslated nucleotides from the beginning of the sequence to the putative start initiation site. The CAPA gene has a single open reading frame, 441 nucleotides long, that codes for a predicted precursor protein of 147 amino acids. The predicted prepropeptide encodes a single copy of each of three deduced propeptides, a CAP2b propeptide, with a Q substituted for an E at the N-terminus (QLYAFPRVa), and two novel CAP2b-related propeptides (DGVLNLYPFPRVa and TEGPGMWFGPRLa). To reduce confusion and to adopt a more standardized nomenclature, we rename pELYAFPRVa as Mas-CAPA-1 and assign the names of Mas-CAPA-2 to DGVLNLYPFPRVa and Mas-PK-1 (Pyrokinin-1) to TEGPGMWFGPRLa. The spatial and temporal expression pattern of the CAPA gene in the Manduca central nervous system (CNS) was determined in all major post-embryonic stages using in situ hybridization techniques. The CAPA gene is expressed in a total of 27 pairs of neurons in the post-embryonic Manduca CNS. A total of 16 pairs of cells is observed in the brain, two pairs in the sub-esophageal ganglion (SEG), one pair in the third thoracic ganglion (T3), one pair in each unfused abdominal ganglion (A1-A6) and two pairs in the fused terminal ganglion. The mRNA from the CAPA gene is present in nearly every ganglion in each post-embryonic stage. The number of cells expressing the CAPA gene varies during post-embryonic life, starting at 54 cells in first-instar larvae and declining to a minimum of 14 cells midway through adult development.

Amino Acid Sequence↗

Random sequencing of cDNA library derived from partially-fed adult female Haemaphysalis longicornis salivary gland.

A cDNA library was constructed from salivary glands of partially-fed adult female Haemaphysalis longicornis (hard tick). Randomly selected clones were sequenced and a total of 633 sequences were analyzed by bioinformatic programs. The sequences were grouped into 213 clusters, with each cluster being considered to be composed of mRNAs derived from the same gene or closely related genes. About 36% of the mRNA sequences showed significant similarity to known proteins in the non-redundant protein database by the NCBI blastx program and appeared to be coding for functional predicted proteins, whereas the remaining 64% had no similar sequences. Two thirds of the predicted proteins were annotated as basic cellular proteins (housekeeping proteins). Among the functional predicted protein sequences, other than the housekeeping proteins, several protease inhibitors including anticoagulants, two metalloproteases and a potential immunosuppressive protein could be identified. These proteins may play important roles during tick feeding and could be novel anti-tick vaccine candidates.

Amino Acid Sequence↗

SNPs on human chromosomes 21 and 22 -- analysis in terms of protein features and pseudogenes.

SNPs are useful for genome-wide mapping and the study of disease genes. Previous studies have focused on SNPs in specific genes or SNPs pooled from a variety of different sources. Here, a systematic approach to the analysis of SNPs in relation to various features on a genome-wide scale, with emphasis on protein features and pseudogenes, is presented. We have performed a comprehensive analysis of 39,408 SNPs on human chromosomes 21 and 22 from the SNP consortium (TSC) database, where SNPs are obtained by random sequencing using consistent and uniform methods. Our study indicates that the occurrence of SNPs is lowest in exons and higher in repeats, introns and pseudogenes. Moreover, in comparing genes and pseudogenes, we find that the SNP density is higher in pseudogenes and the ratio of nonsynonymous to synonymous changes is also much higher. These observations may be explained by the increased rate of SNP accumulation in pseudogenes, which presumably are not under selective pressure. We have also performed secondary structure prediction on all coding regions and found that there is no preferential distribution of SNPs in a -helices, b -sheets or coils. This could imply that protein structures, in general, can tolerate a wide degree of substitutions. Tables relating to our results are available from http://genecensus.org/pseudogene.

Algorithms↗

Complex learning and information processing by pigeons: a critical analysis.

THREE MODELS OF CONDITIONAL DISCRIMINATION LEARNING BY PIGEONS ARE DESCRIBED: stimulus configuration learning, the multiple-rule model, and concept learning. A review of the literature reveals that true concept learning is not characteristic of the behavior of pigeons in matching-to-sample, oddity-from-sample, or symbolic matching studies. Instead, pigeons learn a set of sample-specific S(D) rules. Transfer of the discrimination to novel stimuli, at least along the hue dimension, is predicted by a "coding hypothesis", which holds that pigeons make a unique, but usually unobserved response, R(1), to each sample, and that the comparison stimulus chosen depends on which R(1) was emitted in the presence of the sample. Convincing evidence is found that pigeons do code sample hues, but there is little evidence that allows one to infer that the "coding event" must have behavioral properties. Parameters of the conditional discrimination paradigm are identified, and it is shown that by appropriate parametric manipulation, a variety of analogous tasks may be generated for both human and animal subjects. The tasks make possible the comparative study of complex learning, attention, memory, and information processing, with the added advantage that behavior processes may be compared systematically across tasks.

Journal Article↗

Creating polyketide diversity through genetic engineering.

Modular polyketide synthases (PKS) are large multifunctional enzymes that synthesize complex polyketides, a therapeutically important class of natural products. The linear order and composition of catalytic sites that comprise the PKS represent a "code" that determines the identity of the polyketide product. By re-programming the PKS through genetic engineering, it is possible to alter the code in a predictable manner to create specific structural modifications of polyketides and to produce new libraries of these natural products.

Carbohydrate Sequence↗

An integrated human immunoglobulin germline resource linking allele diversity to expressed repertoire structure.

Human immunoglobulin (IG) loci are highly polymorphic, yet existing germline resources remain noisy and incomplete, limiting our ability to link inherited variation to antibody repertoires. Here, we integrate high-fidelity long-read genomic sequencing with matched adaptive immune receptor repertoire sequencing (AIRR-seq) to construct HUSA, a population-scale, evidence-resolved germline resource. Using a conservative allele inference framework, HUSA expands current references more than three-fold, identifying over 1300 alleles while preserving allele-level evidence provenance across genomic and repertoire data. By linking genotype and expressed repertoires within individuals, we show that coding-region similarity predicts the structure of adjacent recombination signal sequences and leader regions, revealing that IG alleles are organized as linked cis-regulatory units associated with differences in recombination context and allele usage. These results define key germline constraints shaping repertoire formation and establish a robust, genotype-aware foundation for the analysis of immune receptor repertoires.

Journal Article↗

MGC9753 gene, located within PPP1R1B-STARD3-ERBB2-GRB7 amplicon on human chromosome 17q12, encodes the seven-transmembrane receptor with extracellular six-cystein domain.

MYC, ERBB2, MET, FGFR2, CCNE1, MYCN, WNT2, CD44, MDM2, NCOA3, IQGAP1 and STK6 loci are amplified in human gastric cancer. It has been reported that the gene corresponding to EST H16094 is co-amplified with ERBB2 gene in human gastric cancer. Here, we identified and characterized the gene corresponding to EST H16094 by using bioinformatics. BLAST programs revealed that EST H16094 was derived from the uncharacterized MGC9753 gene. Two ORFs were predicted within human MGC9753 mRNA, and ORF1 (nucleotide position 18-980 of NM_033419.1) was predicted as the coding region of human MGC9753 mRNA based on comparative genomics. Nucleotide sequence of mouse Mgc9753 mRNA was next determined in silico by modification of AK052486 cDNA (deleting C at the nucleotide position 37). Human MGC9753 and mouse Mgc9753 proteins were 320-amino-acid seven-transmembrane receptors with the N-terminal six-cysteine domain and an N-glycosylation site (85.0% total-amino-acid identity). Human MGC9753 protein showed 90.6% total-amino-acid identity with human CAB2 aberrant protein, which lacked the third-transmembrane domain of MGC9753 due to frame shifts within ORF. Human MGC9753 gene, consisting of eight exons, were clustered with PPP1R1B, STARD3, TCAP, PNMT, ERBB2, MGC14832 and GRB7 genes within the 120-kb region. PPP1R1B, STARD3, MGC9753, ERBB2 and GRB7 genes are co-amplified in several cases of gastric cancer. This is the first report on comprehensive characterization of the amplicon around the PPP1R1B-STARD3-TCAP-PNMT-MGC9753-ERBB2-MGC14832-GRB7 locus on human chromosome 17q12.

Amino Acid Sequence↗

Molecular cloning of a homeobox transcription factor from adult aortic smooth muscle.

We report here the cloning of a cDNA encoding a homeobox transcription factor from vascular smooth muscle and describe its unique pattern of mRNA expression at different stages in development. The cDNA isolated is 1576 base pairs in length not including the poly(A) tail and contains an open reading frame coding for a predicted 372-amino acid homeobox protein. During early embryogenesis, expression was detected in the neural tube with a sharp expression boundary occurring at an anterior position, in the myelencephalon, in the third and fourth branchial arches, and in vessels leading from the heart. In adults, however, transcripts were only detected in aortic smooth muscle and lung but were undetectable in cardiac or skeletal muscle, visceral smooth muscle, and many other tissues including brain. In neonates, expression was detected in the outflow tracts of the heart as well as in the cardiomyocytes. The expression pattern of this gene suggests that, although it likely has multiple roles during development, in the adult, it may participate in the control of vascular smooth muscle differentiation and proliferation.

Aging↗

Identification and characterization of human KIAA1391 and mouse Kiaa1391 genes encoding novel RhoGAP family proteins with RA domain and ANXL repeats.

The commonly deleted region of breast cancer at human chromosome 11q23.1 is located between microsatellite markers D11S927 and D11S1347. We have previously identified and characterized the KIAA1735 gene, encoding MTH and DIX domain protein, within the 11q23.1 commonly deleted region. The BTG4 gene, encoding BTG/TOB family protein, as well as the SIK2 gene, encoding salt-inducible serine/threonine kinase 2, were also located within the 11q23.1 region. We identified and characterized the KIAA1391 gene within the 11q23.1 commonly deleted region by using bioinformatics. The nucleotide position 11-3586 of human KIAA1391 cDNA was predicted as the coding region of human KIAA1391 gene. Mouse Kiaa1391 gene was located within mouse genome sequences RP23-95E13 (AC102770.4), RP23-212M11 (AC140363.1), and mouse chromosome 9 genomic contig NT_039473.1. Exon-intron structure was well conserved between human KIAA1391 gene and mouse Kiaa1391 gene. Nucleotide sequence of mouse Kiaa1391 cDNA was determined in silico by assembling nucleotide sequences of exon 1-15 of mouse Kiaa1391 gene. Human KIAA1391 protein (1191 aa) and mouse Kiaa1391 protein (1182 aa) showed 79.1% total-amino-acid identity. RA domain (codon 198-274), RhoGAP domain (codon 363-546), and two Annexin-like (ANXL) repeats (codon 657-767 and 796-909) of human KIAA1391 protein were conserved in mouse Kiaa1391 protein. This is the first report on the comprehensive characterization of human KIAA1391 gene and mouse Kiaa1391 gene, encoding RhoGAP proteins with RA domain and two ANXL repeats.

Amino Acid Sequence↗

Order sets utilization in a clinical order entry system.

An order set is a predefined template that has been utilized in the standard care of hospitals for many years. While in the past, it took the form of pen and paper, today, it is, indeed, electronic. Within order sets are distinct ordering patterns that may yield fruitful results for clinicians and informaticians, alike. Protocols like there electronic counterpart, order sets, provide an 'indication' identifying the clinical scenario of the patient's condition when the ordering event occurred. This 'indication' is rarely captured by individual orders, and provides difficult challenges to developers of information systems. While mandating an 'indication' be entered for every medication or lab order makes the job much more tasking on the physician provider, it is appealing to researchers and accountants. We have attempted to bypasses that consideration by identifying ordering patterns that predict diagnostic related codes (DRGs) and diagnostic codes which would greatly facilitate the information gathering process and still provide a flexible and user friendly physician interface.

Forms and Records Control↗