Search PubMedSearch

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Functional expression cloning of the canalicular sulfate transport system of rat hepatocytes.

We have cloned a single cDNA encoding the canalicular sulfate transporter of rat liver using Xenopus laevis oocytes as a functional expression system. The cloned cDNA sulfate anion transporter-1 (sat-1) expresses saturable Na(+)-independent sulfate uptake (Km approximately 0.14 mM) that can be inhibited by 4,4'-diisothiocyano-2,2'-disulfonic acid stilbene (DIDS, IC50 = 28 microM) and oxalate, but not by succinate or cholate. These properties are very similar to sulfate uptake expressed in oocytes injected with total rat liver mRNA and to the bicarbonate/sulfate exchange system previously characterized in canalicular rat liver plasma membrane vesicles. The cloned sat-1 cDNA has a total length of 3726 base pairs (bp) with an open reading frame encompassing 2109 bp, a 5'-untranslated region of 367 bp, and a 3'-untranslated region of 1250 bp. The coding region predicts a protein of 703 amino acids with a calculated molecular mass of 75.4 kDa. Computer-based hydrophobicity analysis suggests the presence of 12 putative transmembrane spanning domains. Furthermore, three potential glycosylation sites are detected (Asn-158, Asn-163, Asn-587). Northern blot analysis indicates that similar sulfate anion transporters are also present in the kidney, muscle, and brain of rat and in the liver of the mouse. Using antisense oligonucleotides the mRNA-species of the sat-1 analogue in rat kidney has been characterized by hybrid depletion experiments (Markovich, D., Bissig, M., Sorribas, V., Hagenbuch, B., Meier, P. J., and Murer, H. (1994) J. Biol. Chem. 269, 3022-3026).

Amino Acid Sequence

The incidence of herpes zoster.

BACKGROUND: There are few population-based studies of the natural history and epidemiology of herpes zoster. Although a relatively common cause of morbidity, especially among the elderly, contemporary estimates of herpes zoster incidence are lacking. Herein we describe a population-based investigation of incident and recurrent herpes zoster from 1990 through 1992 in a health maintenance organization. METHODS: The health maintenance organization's automated medical records contain clinical and administrative information about care rendered to patients in ambulatory settings, emergency departments, and hospitals. Cases of herpes zoster were ascertained by screening the medical record for coded diagnoses. The predictive value of a herpes zoster diagnosis code was determined by review of a sample of patient records. Records from all patients with potential recurrences were also reviewed. RESULTS: The overall incidence, based on 1075 cases in 500,408 person-years, was 215 per 100,000 person-years (95% confidence interval, 192 to 240 per 100,000) and did not vary by gender. Although the rate increased sharply with age, approximately 5% of the cases occurred among children younger than 15 years. Infection with human immunodeficiency virus was documented in 5% of the persons with incident herpes zoster and cancer in 6%. Four persons had confirmed recurrences of herpes zoster (744 per 100,000 person-years; 95% confidence interval, 203 to 1907); three of these persons were infected with the human immunodeficiency virus. CONCLUSIONS: The recorded incidence of herpes zoster was 64% higher than that reported 30 years ago; the age-standardized rate was more than twofold higher. Immunosuppressive conditions had little impact on overall incidence, although they were strongly associated with early recurrences.

Adolescent

Extensive sequence homology of the goldfish ras gene to mammalian ras genes.

We cloned ras-related sequences from goldfish genomic libraries constructed as recombinants using the lambda phage. Restriction enzyme mapping of the clones obtained revealed three kinds of ras-related sequences among approximately 350,000 genomic clones. One of these clones was partially sequenced. Comparison with the nucleotide sequences of mammalian ras genes showed that the determined sequences covered the predicted amino acid coding regions and parts of the intervening regions. The predicted amino acid sequences of the cloned ras-related goldfish gene suggested that the coding region is localized separately in DNA, and that its exon-intron boundaries are exactly the same as those of corresponding mammalian genes. The nucleotide and amino acid sequences of the goldfish ras-related gene may have extensive homologies to mammalian p 21 protein. Among the three mammalian ras proteins, the predicted amino acid sequence of the sequenced ras-related goldfish clone is most closely homologous (96%) to the Kirsten ras protein. Differences in the predicted amino acid sequence were greatest in the sequence predicted from the fourth exon; fewer differences were found in the sequence from the third exon, and only slight or no differences were found in the sequence predicted for the first and second exons. The 12th and 61st amino acids from the N-terminal of the protein, which are thought to be critical positions for GTP binding and catalysis, are both conserved in the goldfish protein.(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence

Identification of coding regions in genomic DNA sequences: an application of dynamic programming and neural networks.

Dynamic programming (DP) is applied to the problem of precisely identifying internal exons and introns in genomic DNA sequences. The program GeneParser first scores the sequence of interest for splice sites and for these intron- and exon-specific content measures: codon usage, local compositional complexity, 6-tuple frequency, length distribution and periodic asymmetry. This information is then organized for interpretation by DP. GeneParser employs the DP algorithm to enforce the constraints that introns and exons must be adjacent and non-overlapping and finds the highest scoring combination of introns and exons subject to these constraints. Weights for the various classification procedures are determined by training a simple feed-forward neural network to maximize the number of correct predictions. In a pilot study, the system has been trained on a set of 56 human gene fragments containing 150 internal exons in a total of 158,691 bps of genomic sequence. When tested against the training data, GeneParser precisely identifies 75% of the exons and correctly predicts 86% of coding nucleotides as coding while only 13% of non-exon bps were predicted to be coding. This corresponds to a correlation coefficient for exon prediction of 0.85. Because of the simplicity of the network weighting scheme, generalization performance is nearly as good as with the training set.

Algorithms

Cloning and sequencing the HinfI restriction and modification genes.

The HinfI restriction and modification genes were cloned on a 3.9-kb PstI fragment inserted into the PstI site of plasmid pBR322. Both genes are confined to an internal 2.3-kb BclI-AvaI subfragment. This subfragment was sequenced. Two large open reading frames (ORF's) are present. ORF1 codes for the methylase [predicted 359 amino acids (aa)] and ORF2 codes for the endonuclease (predicted 262 or 272 aa).

Amino Acid Sequence

Identification of human gene structure using linear discriminant functions and dynamic programming.

Development of advanced technique to identify gene structure is one of the main challenges of the Human Genome Project. Discriminant analysis was applied to the construction of recognition functions for various components of gene structure. Linear discriminant functions for splice sites, 5'-coding, internal exon, and 3'-coding region recognition have been developed. A gene structure prediction system FGENE has been developed based on the exon recognition functions. We compute a graph of mutual compatibility of different exons and present a gene structure models as paths of this directed acyclic graph. For an optimal model selection we apply a variant of dynamic programming algorithm to search for the path in the graph with the maximal value of the corresponding discriminant functions. Prediction by FGENE for 185 complete human gene sequences has 81% exact exon recognition accuracy and 91% accuracy at the level of individual exon nucleotides with the correlation coefficient (C) equals 0.90. Testing FGENE on 35 genes not used in the development of discriminant functions shows 71% accuracy of exact exon prediction and 89% at the nucleotide level (C = 0.86). FGENE compares very favorably with the other programs currently used to predict protein-coding regions. Analysis of uncharacterized human sequences based on our methods for splice site (HSPL, RNASPL), internal exons (HEXON), all type of exons (FEXH) and human (FGENEH) and bacterial (CDSB) gene structure prediction and recognition of human and bacterial sequences (HBR) (to test a library for E. coli contamination) is available through the University of Houston, Weizmann Institute of Science network server and a WWW page of the Human Genome Center at Baylor College of Medicine.

Algorithms

GraphyloVar: predicting the impact of non-coding variants using a multi-species sequence model.

MOTIVATION: Understanding the functional impact of genetic variants is a key problem for precision medicine. Tools like CADD, PhyloP, and PhastCons are useful, but they often look at each position in the genome in isolation. This means they can miss important information from the evolutionary history that connects different species. In this paper, we extend our previous model, Graphylo, to predict the effects of variants. Our new model, GraphyloVar, is built to directly utilize the phylogenetic tree that relates the species. RESULTS: GraphyloVar is a deep learning model that considers both DNA sequence and evolutionary patterns from many species. It uses two main components: Graph Convolutional Networks (GCNs) to process the phylogenetic tree, and Transformer encoders to extract features from the DNA sequences. Pre-trained to predict population-level allele frequencies on the TOPMed whole-genome sequencing cohort, GraphyloVar achieves an AUROC of 0.6246 zero-shot on &#x223c;149M held-out variants, and an ensemble with CADD reaches 0.6442 (+0.020, P<10-15). Fine-tuned GraphyloVar achieves the highest AUROC across all 13 MPRA benchmark datasets. By integrating deep learning with explicit phylogenetic input, GraphyloVar offers a powerful and complementary approach to variant effect prediction that utilizes the full evolutionary history from many species to better identify and prioritize important non-coding variants. AVAILABILITY AND IMPLEMENTATION: Code and datasets are available at https://github.com/DongjoonLim/GraphyloVar under DOI: 10.5281/zenodo.20616818.

Phylogeny

Chromosome-level genome assembly of Nothapodytes nimmoniana.

Nothapodytes nimmoniana is a plant species belonging to the genus Nothapodytes in the family Icacinaceae. This species holds significant medicinal value due to its camptothecin content. In this study, we present the first chromosome-level genome assembly of N. nimmoniana constructed using NGS, Hi-C, and HiFi sequencing technologies. The assembled genome spans 3.53&#x2009;Gb across 14 chromosomes, with an N50 length of 248.74&#x2009;Mb. Genome annotation revealed that repetitive sequences constitute 80.82% of the genome size, and 83,269 protein-coding genes were predicted. Additionally, 4,360,538&#x2009;bp of non-coding RNA were annotated. This genomic resource provides a foundation for further investigation into camptothecin biosynthesis pathways and plant phylogeny in N. nimmoniana.

Genome, Plant

Concordance of experimentally mapped or predicted Z-DNA sites with positions of selected alternating purine-pyrimidine tracts.

The recent electronmicroscopic and biochemical mapping of Z-DNA sites in phi X174, SV40, pBR322 and PM2 DNAs has been used to determine two sets of criteria for identification of potential Z-DNA sequences in natural DNA genomes. The prediction of potential Z-DNA tracts and corresponding statistical analysis of their occurrence have been made on a sample of 14 DNA genomes. Alternating purine and pyrimidine tracts longer than 5 base pairs in length and their clusters (quasi alternating fragments) in the 14 genomes studied are under-represented compared to the expectation from corresponding random sequences. The fragments [d(G X C)]n and [d(C X G)]n (n greater than or equal to 3) in general do not occur in circular DNA genomes and are under-represented in the linear DNAs of phages lambda and T7, whereas in linear genomes of adenoviruses they are strongly over-represented. With minor exceptions, potential Z-DNA sites are also under-represented compared to random sequences. In the 14 genomes studied, predicted Z-DNA tracts occur in non-coding as well as in protein coding regions. The predicted Z-DNA sites in phi X174, SV40, pBR322 and PM2 correspond well with those mapped experimentally. A complete listing together with a compact graphical representation of alternating purine-pyrimidine fragments and their Z-forming potential are presented.

Animals

Isolation and characterization of a partial cDNA for a human sialyltransferase.

A probe generated from the coding sequence of the rat hepatic beta-galactoside alpha 2,6-sialyltransferase was used to screen a human cDNA library constructed of human submaxillary gland mRNA lambda gt-11. We report the isolation and characterization of a human cDNA, HSM-ST1, that is putatively the human homolog of the beta-galactoside alpha 2,6-sialyltransferase. The largest human clone contains a 1.3 kb cDNA insert and is predicted to encompass 75% of the coding sequence as well as a small portion of the 3' untranslated region. Comparative analysis of this insert with the rat hepatic alpha 2,6-sialyltransferase sequence indicates 79% nucleotide similarity between the two sequences in the predicted coding region. On the amino acid level, the degree of conservation is 86%. Substantial sequence similarity is observed in the 3'-untranslated region between the rat and human sequences as well. S1 nuclease analysis was performed to demonstrate the expression of HSM-ST1 transcripts in the human hepatoma cell line, HepG2, and in the human colonic adenocarcinoma cell lines, LS174T.

Amino Acid Sequence

Chromosomal level genome assembly of medicinal plant Chrysosplenium macrophyllum.

Chrysosplenium macrophyllum Oliv., a perennial herb native to China, is widely used in traditional medicine for its notable therapeutic properties. However, the absence of a reference genome has constrained its full potential for research and application. This study presents the first chromosome-level de novo genome assembly of C. macrophyllum, constructed by integrating long reads from Oxford Nanopore Technologies (ONT), short reads from BGI, and Hi-C data. The final assembly spans 2.55&#x2009;Gb, with a scaffold N50 of 93.38&#x2009;Mb, and 83.70% of the genome has been assigned to 22 chromosomes. The mapping rate of the BGI short reads to the genome is approximately 97.94%, and BUSCO analysis reveals that 97.94% of the predicted genes are complete. A total of 62,921 protein-coding genes were predicted, with functional annotations for 93.67% of them. This chromosome-level genome assembly represents an important resource for expanding our understanding of Chrysosplenium species and supports future genomic studies and applications.

Genome, Plant

Variable deletion of exon 9 coding sequences in cystic fibrosis transmembrane conductance regulator gene mRNA transcripts in normal bronchial epithelium.

The predicted protein domains coded by exons 9-12 and 19-23 of the 27 exon cystic fibrosis transmembrane conductance regulator (CFTR) gene contain two putative nucleotide-binding fold regions. Analysis of CFTR mRNA transcripts in freshly isolated bronchial epithelium from 12 normal adult individuals demonstrated that all had some CFTR mRNA transcripts with exon 9 completely deleted (exon 9- mRNA transcripts). In most (9 of 12), the exon 9- transcripts represented less than or equal to 25% of the total CFTR transcripts. However, in three individuals, the exon 9- transcripts were more abundant, comprising 39, 62 and 66% of all CFTR transcripts. Re-evaluation of the same individuals 2-4 months later showed the same proportions of exon 9- transcripts. Of the 24 CFTR alleles in the 12 individuals, the sequences of the exon-intron junctions relevant to exon 9 deletion (exon 8-intron 8, intron 8-exon 9, exon 9-intron 9, and intron 9-exon 10) were identical except for the intron 8-exon 9 region sequences. Several individuals had varying lengths of a TG repeat in the region between splice branch and splice acceptor consensus sites. Interestingly, one allele in each of the two individuals with 62 and 66% exon 9- transcripts had a TT deletion in the splice acceptor site for exon 9. These observations suggest either the unlikely possibility that sequences in exon 9 are not critical for the functioning of the CFTR or that only a minority of the CFTR mRNA transcripts need to contain exon 9 sequences to produce sufficient amounts of a normal CFTR to maintain a normal clinical phenotype.

Base Sequence

Correlation approach to identify coding regions in DNA sequences.

Recently, it was observed that noncoding regions of DNA sequences possess long-range power-law correlations, whereas coding regions typically display only short-range correlations. We develop an algorithm based on this finding that enables investigators to perform a statistical analysis on long DNA sequences to locate possible coding regions. The algorithm is particularly successful in predicting the location of lengthy coding regions. For example, for the complete genome of yeast chromosome III (315,344 nucleotides), at least 82% of the predictions correspond to putative coding regions; the algorithm correctly identified all coding regions larger than 3000 nucleotides, 92% of coding regions between 2000 and 3000 nucleotides long, and 79% of coding regions between 1000 and 2000 nucleotides. The predictive ability of this new algorithm supports the claim that there is a fundamental difference in the correlation property between coding and noncoding sequences. This algorithm, which is not species-dependent, can be implemented with other techniques for rapidly and accurately locating relatively long coding regions in genomic sequences.

Algorithms

Nucleotide sequence of the pilin gene of Bacteroides nodosus 340 (serogroup D) and implications for the relatedness of serogroups.

The gene encoding pilin of Bacteroides nodosus 340 has been isolated and the nucleotide sequence determined. The gene is present as a single copy within the B. nodosus genome and a protein of Mr 16683 can be predicted from the proposed coding region. A comparison of the predicted amino acid sequence with pilin from other strains of B. nodosus indicated that the protein of strain 340 (serogroup D) has a high degree of similarity with pilin of strain 265 (serogroup H). The degree of similarity between pilins from these strains and from other B. nodosus serogroups is no greater than that between B. nodosus pilins and the homologous proteins of several different bacterial species. These findings suggest that serogroups D and H may form a subset of B. nodosus serogroups.

Amino Acid Sequence

Cloning and nucleotide sequence analysis of the dog insulin gene. Coded amino acid sequence of canine preproinsulin predicts an additional C-peptide fragment.

A 4.0-kilobase HindIII/EcoRI-cleaved dog genomic DNA fragment was shown to contain the dog insulin gene by restriction mapping using a human insulin cDNA probe. This fragment was subsequently cloned in a lambda vector, and the nucleotide sequence of the dog insulin gene was determined. As in several other species, the insulin gene of the dog is interrupted by two intervening sequences, one of 151 base pairs located in the 5' untranslated region and the other of 264 base pairs occurring within the codon of the 7th amino acid of the C-peptide. Translation of the nucleotide sequence in one frame revealed the primary structure of canine preproinsulin. An interesting feature of the coded amino acid sequence is that it predicts a C-peptide of 31 amino acids, 8 residues longer than that reported by Peterson et al. (Peterson, J. D., Nehrlich, S., Oyer, P. E., and Steiner, D. F. (1973) J. Biol. Chem. 247, 4866-4871). The additional octapeptide sequence, Glu-Val-Glu-Asp-Leu-Gln-Val-Arg, is located NH2-terminal to the 23-residue C-peptide sequence described in the earlier report. Its coding sequence is interrupted by the second intervening sequence. The arginine at position 8 suggests that a trypsin-like cleavage may separate the NH2-terminal octapeptide from the remainder of the C-peptide during the post-translational processing of dog proinsulin in the pancreas. The revised C-peptide sequence suggests that the proinsulin C-peptide is more highly conserved in length and overall sequence than was previously supposed.

Amino Acid Sequence

The hemoglobin of Urechis caupo. The cDNA-derived amino acid sequence.

The nucleotide sequence of a cDNA transcript containing part of the 5' noncoding region, the entire coding region, and the entire 3' noncoding region has been determined. The protein sequence predicted from the coding region matches almost exactly the aminoterminal sequence and the sequence of several peptides from Urechis caupo F-I globin. Only 11-20% of the amino acid positions are identical with those of other known globins.

Amino Acid Sequence

Leveraging functional annotations to map rare variants associated with Alzheimer disease with gruyere.

Increased availability of whole-genome sequencing (WGS) has facilitated the study of rare variants (RVs) in complex diseases. Multiple RV association tests are available to study the relationship between genotype and phenotype, but most do not fully leverage the availability of variant-level functional annotations. We propose genome-wide rare variant enrichment evaluation (gruyere), an empirical Bayesian framework that complements existing methods by learning global, trait-specific weights for functional annotations to improve variant prioritization. We apply gruyere to WGS data from the Alzheimer's Disease Sequencing Project to identify Alzheimer disease (AD)-associated genes and annotations. Growing evidence suggests that the disruption of microglial regulation is a key contributor to AD risk, yet existing methods have not examined rare non-coding effects that incorporate such cell-type-specific information. To address this gap, we (1) define per-gene non-coding RV test sets using predicted enhancer and promoter regions in microglia and other brain cell types (oligodendrocytes, astrocytes, and neurons) and (2) include cell-type-specific variant effect predictions (VEPs) as functional annotations. gruyere identifies 13 significant genetic associations not detected by other RV methods, four of which remain significant in omnibus tests. We find that deep-learning-based VEPs for splicing, transcription factor binding, and chromatin state are highly predictive of functional non-coding RVs. Our study establishes a robust framework incorporating functional annotations, coding RVs, and cell-type-associated non-coding RVs to perform genome-wide association tests, uncovering AD-relevant genes and annotations.

Alzheimer Disease

A 10,400-molecular-weight membrane protein is coded by region E3 of adenovirus.

Previous studies with adenovirus mutants have indicated that a 10,400-molecular-weight (10.4K) protein predicted to be coded by an open reading frame in region E3 of adenovirus functions to down regulate the epidermal growth factor receptor (C. R. Carlin, A. E. Tollefson, H. A. Brady, B. L. Hoffman, and W. S. M. Wold, Cell 57:135-144, 1989). We now demonstrate that the 10.4K protein is in fact synthesized in cells infected by group C adenoviruses. This was done by immunoprecipitation of 10.4K from cells infected by a variety of E3 mutants, using antisera against three different synthetic peptides corresponding to the predicted 10.4K sequence. The 10.4K protein was translated primarily from E3 mRNA f, as indicated by cell-free translation of mRNA purified by hybridization from cells infected with an RNA processing mutant that synthesizes predominantly mRNA f. The 10.4K protein was overproduced or underproduced in vivo, respectively, by mutants that overproduce or underproduce E3 mRNA f, also indicating that the 10.4K protein is translated primarily from mRNA f. The 10.4K protein migrated as two bands with apparent molecular weights of 16,000 and 11,000 (10 to 18% gradient gels); both bands contained 10.4K epitopes, as shown by Western blot (immunoblot). Only the 16K band was obtained by cell-free translation, suggesting that the 16K protein is the precursor to the 11K protein. The 10.4K protein is a membrane protein, as shown by cell fractionation experiments and as predicted from its sequence. The predicted 10.4K sequence as well as a putative N-terminal signal sequence and 30-residue transmembrane domain are conserved in adenovirus types 2 and 5 (group C) and in types 3, 7, and 35 (group B).

Adenoviruses, Human