Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Isolation and molecular evolution of the selenocysteine tRNA (Cf TRSP) and RNase P RNA (Cf RPPH1) genes in the dog family, Canidae.

In an effort to identify rapidly evolving nuclear sequences useful for phylogenetic analyses of closely related species, we isolated two genes transcribed by RNA polymerase III (pol III), the selenocysteine tRNA gene (TRSP) and an RNase P RNA (RPPH1) gene from the domestic dog (Canis familiaris). We focus on genes transcribed by pol III because their coding regions are small (generally 100-300 base pairs [bp]) and their essential promoter elements are located within a couple of hundred bps upstream of the coding region. Therefore, we predicted that regions flanking the coding region and outside of the promoter elements would be free of constraint and would evolve rapidly. We amplified TRSP from 23 canids and RPPH1 from 12 canids and analyzed the molecular evolution of these genes and their utility as phylogenetic markers for resolving relationships among species in Canidae. We compared the rate of evolution of the gene-flanking regions to other noncoding regions of nuclear DNA (introns) and to the mitochondrial encoded COII gene. Alignment of TRSP from 23 canids revealed that regions directly adjacent to the coding region display high sequence variability. We discuss this pattern in terms of functional mechanisms of transcription. Although the flanking regions evolve no faster than introns, both genes were found to be useful phylogenetic markers, in part, because of the synapomorphic indels found in the flanking regions. Gene trees generated from the TRSP and RPPH1 loci were generally in agreement with the published mtDNA phylogeny and are the first phylogeny of Canidae based on nuclear sequences.

Animals↗

Comprehensive analysis of alternative splicing in rice and comparative analyses with Arabidopsis.

BACKGROUND: Recently, genomic sequencing efforts were finished for Oryza sativa (cultivated rice) and Arabidopsis thaliana (Arabidopsis). Additionally, these two plant species have extensive cDNA and expressed sequence tag (EST) libraries. We employed the Program to Assemble Spliced Alignments (PASA) to identify and analyze alternatively spliced isoforms in both species. RESULTS: A comprehensive analysis of alternative splicing was performed in rice that started with >1.1 million publicly available spliced ESTs and over 30,000 full length cDNAs in conjunction with the newly enhanced PASA software. A parallel analysis was performed with Arabidopsis to compare and ascertain potential differences between monocots and dicots. Alternative splicing is a widespread phenomenon (observed in greater than 30% of the loci with transcript support) and we have described nine alternative splicing variations. While alternative splicing has the potential to create many RNA isoforms from a single locus, the majority of loci generate only two or three isoforms and transcript support indicates that these isoforms are generally not rare events. For the alternate donor (AD) and acceptor (AA) classes, the distance between the splice sites for the majority of events was found to be less than 50 basepairs (bp). In both species, the most frequent distance between AA is 3 bp, consistent with reports in mammalian systems. Conversely, the most frequent distance between AD is 4 bp in both plant species, as previously observed in mouse. Most alternative splicing variations are localized to the protein coding sequence and are predicted to significantly alter the coding sequence. CONCLUSION: Alternative splicing is widespread in both rice and Arabidopsis and these species share many common features. Interestingly, alternative splicing may play a role beyond creating novel combinations of transcripts that expand the proteome. Many isoforms will presumably have negative consequences for protein structure and function, suggesting that their biological role involves post-transcriptional regulation of gene expression.

Alternative Splicing↗

Postgenomic bioinformatic analysis of yeast artificial chromosome sequence.

The free availability of multiple genomic sequences represents one of the greatest advances in biology of the new millennium, and promises to revolutionize our ability to determine and treat the causes of human disease. This chapter highlights a number of basic, freely available, and user-friendly bioinformatic techniques that can be used to predict the functional genetic contents of specific yeast artificial chromosome (YAC) clones. The content of this chapter is written for the level of graduate students, who may be relatively inexperienced with the use of computers for analyzing DNA sequences. The basic instructions that allow the identification of the genomic sequence of interest and to download this sequence onto a personal computer from an online database are presented. Simple instructions are also given on how to perform basic sequence manipulations, how to use online tools to design polymerase chain reaction primers, and how to map restriction sites. Also described are more complicated programs that rapidly and efficiently perform genome alignments that, in addition to predicting the location of protein coding sequences, allow the prediction of functional genomic sequences, such as cis regulatory elements and scaffold/matrix attachment sites. The availability of genomic sequences and the rapidly expanding numbers of predictive programs that allow the predictive analysis of these sequences promises to greatly facilitate the use of YAC clones in the search for the causes of disease.

Chromosomes, Artificial, Yeast↗

Sequence and analysis of chromosome 3 of the plant Arabidopsis thaliana.

Arabidopsis thaliana is an important model system for plant biologists. In 1996 an international collaboration (the Arabidopsis Genome Initiative) was formed to sequence the whole genome of Arabidopsis and in 1999 the sequence of the first two chromosomes was reported. The sequence of the last three chromosomes and an analysis of the whole genome are reported in this issue. Here we present the sequence of chromosome 3, organized into four sequence segments (contigs). The two largest (13.5 and 9.2 Mb) correspond to the top (long) and the bottom (short) arms of chromosome 3, and the two small contigs are located in the genetically defined centromere. This chromosome encodes 5,220 of the roughly 25,500 predicted protein-coding genes in the genome. About 20% of the predicted proteins have significant homology to proteins in eukaryotic genomes for which the complete sequence is available, pointing to important conserved cellular functions among eukaryotes.

Arabidopsis↗

Chromosome-level genome assembly of Nothapodytes nimmoniana.

Nothapodytes nimmoniana is a plant species belonging to the genus Nothapodytes in the family Icacinaceae. This species holds significant medicinal value due to its camptothecin content. In this study, we present the first chromosome-level genome assembly of N. nimmoniana constructed using NGS, Hi-C, and HiFi sequencing technologies. The assembled genome spans 3.53 Gb across 14 chromosomes, with an N50 length of 248.74 Mb. Genome annotation revealed that repetitive sequences constitute 80.82% of the genome size, and 83,269 protein-coding genes were predicted. Additionally, 4,360,538 bp of non-coding RNA were annotated. This genomic resource provides a foundation for further investigation into camptothecin biosynthesis pathways and plant phylogeny in N. nimmoniana.

Genome, Plant↗

Concordance of experimentally mapped or predicted Z-DNA sites with positions of selected alternating purine-pyrimidine tracts.

The recent electronmicroscopic and biochemical mapping of Z-DNA sites in phi X174, SV40, pBR322 and PM2 DNAs has been used to determine two sets of criteria for identification of potential Z-DNA sequences in natural DNA genomes. The prediction of potential Z-DNA tracts and corresponding statistical analysis of their occurrence have been made on a sample of 14 DNA genomes. Alternating purine and pyrimidine tracts longer than 5 base pairs in length and their clusters (quasi alternating fragments) in the 14 genomes studied are under-represented compared to the expectation from corresponding random sequences. The fragments [d(G X C)]n and [d(C X G)]n (n greater than or equal to 3) in general do not occur in circular DNA genomes and are under-represented in the linear DNAs of phages lambda and T7, whereas in linear genomes of adenoviruses they are strongly over-represented. With minor exceptions, potential Z-DNA sites are also under-represented compared to random sequences. In the 14 genomes studied, predicted Z-DNA tracts occur in non-coding as well as in protein coding regions. The predicted Z-DNA sites in phi X174, SV40, pBR322 and PM2 correspond well with those mapped experimentally. A complete listing together with a compact graphical representation of alternating purine-pyrimidine fragments and their Z-forming potential are presented.

Animals↗

Sequence and organization of pMAC, an Acinetobacter baumannii plasmid harboring genes involved in organic peroxide resistance.

Acinetobacter baumannii 19606 harbors pMAC, a 9540-bp plasmid that contains 11 predicted open-reading frames (ORFs). Cloning and transformation experiments using Acinetobacter calcoaceticus BD413 mapped replication functions within a region containing four 21-bp direct repeats (ori) and ORF 1, which codes for a predicted replication protein. Subcloning and tri-parental mating experiments mapped mobilization functions to the product of ORF 11 and an adjacent predicted oriT. Three ORFs code for proteins that share similarity to hypothetical proteins encoded by plasmid genes found in other bacteria, while the predicted products of three others do not match any known sequence. The product of ORF 8 is similar to Ohr, a hydroperoxide reductase responsible for organic peroxide detoxification and resistance in bacteria. This ORF is immediately upstream of a coding region whose product is related to the MarR family of transcriptional regulators. Disk diffusion assays showed that A. baumannii 19606 is resistant to the organic peroxide-generating compounds cumene hydroperoxide (CHP) and tert-butyl hydroperoxide (t-BHP), although to levels lower than those detected in Pseudomonas aeruginosa PAO1. Cloning and introduction of the ohr and marR ORFs into Escherichia coli was associated with an increase in resistance to CHP and t-BHP. This appears to be the first case in which the genetic determinants involved in organic peroxide resistance are located in an extrachromosomal element, a situation that can facilitate the horizontal transfer of genetic elements coding for a function that protects bacterial cells from oxidative damage.

Acinetobacter baumannii↗

Complete genome sequence of the industrial bacterium Bacillus licheniformis and comparisons with closely related Bacillus species.

BACKGROUND: Bacillus licheniformis is a Gram-positive, spore-forming soil bacterium that is used in the biotechnology industry to manufacture enzymes, antibiotics, biochemicals and consumer products. This species is closely related to the well studied model organism Bacillus subtilis, and produces an assortment of extracellular enzymes that may contribute to nutrient cycling in nature. RESULTS: We determined the complete nucleotide sequence of the B. licheniformis ATCC 14580 genome which comprises a circular chromosome of 4,222,336 base-pairs (bp) containing 4,208 predicted protein-coding genes with an average size of 873 bp, seven rRNA operons, and 72 tRNA genes. The B. licheniformis chromosome contains large regions that are colinear with the genomes of B. subtilis and Bacillus halodurans, and approximately 80% of the predicted B. licheniformis coding sequences have B. subtilis orthologs. CONCLUSIONS: Despite the unmistakable organizational similarities between the B. licheniformis and B. subtilis genomes, there are notable differences in the numbers and locations of prophages, transposable elements and a number of extracellular enzymes and secondary metabolic pathway operons that distinguish these species. Differences include a region of more than 80 kilobases (kb) that comprises a cluster of polyketide synthase genes and a second operon of 38 kb encoding plipastatin synthase enzymes that are absent in the B. licheniformis genome. The availability of a completed genome sequence for B. licheniformis should facilitate the design and construction of improved industrial strains and allow for comparative genomics and evolutionary studies within this group of Bacillaceae.

Anti-Bacterial Agents↗

Isolation and characterization of a partial cDNA for a human sialyltransferase.

A probe generated from the coding sequence of the rat hepatic beta-galactoside alpha 2,6-sialyltransferase was used to screen a human cDNA library constructed of human submaxillary gland mRNA lambda gt-11. We report the isolation and characterization of a human cDNA, HSM-ST1, that is putatively the human homolog of the beta-galactoside alpha 2,6-sialyltransferase. The largest human clone contains a 1.3 kb cDNA insert and is predicted to encompass 75% of the coding sequence as well as a small portion of the 3' untranslated region. Comparative analysis of this insert with the rat hepatic alpha 2,6-sialyltransferase sequence indicates 79% nucleotide similarity between the two sequences in the predicted coding region. On the amino acid level, the degree of conservation is 86%. Substantial sequence similarity is observed in the 3'-untranslated region between the rat and human sequences as well. S1 nuclease analysis was performed to demonstrate the expression of HSM-ST1 transcripts in the human hepatoma cell line, HepG2, and in the human colonic adenocarcinoma cell lines, LS174T.

Amino Acid Sequence↗

Evaluation of a (1)h-(13)c NMR spectral library.

A simple database of (13)C/(1)H-(13)C spectral lists for 11 673 natural products was created in standard commercial database format. Over 50% of the spectra were predicted using HOSE code descriptors derived from the 50% of spectra having experimental values. Prediction errors obtained by prediction of and comparison to the experimental spectra revealed an exponentially decaying dependence between the average absolute error and the depth of the matching HOSE codes. A subset of the library containing over 1000 (1)H-(13)C assigned experimental spectral lists were used to test against eight alternate query data sets. These sets represent query data from various combinations of 1D-(13)C, 1D-DEPT, and 2D-(1)H-(13)C spectra. Simulated query lists were generated using Monte Carlo methods. As expected, queries based on 2D-(1)H-(13)C data were more likely to find the correct match under unfavorable conditions.

Journal Article↗

Solid-solution metal partitioning in the Humber rivers: application of WHAM and SCAMP

The solid-solution partitioning of five trace metals (Co, Ni, Cu, Zn and Pb) in four LOIS rivers has been modelled using a predictive chemical speciation code (WHAM-SCAMP). Observed (log K(D,obs)) and predicted (log K(D,pred)) log K(D) values were similar for Zn. For Co and Ni, the log K(D,pred) values were typically greater than the log K(D,obs) values, while for Cu and Pb the reverse was seen. Removing modelled competition by Ca and Mg for binding sites on particulate organic matter increased the log K(D,pred) values for all the metals except Pb, and gave better agreement between the observations and the predictions for Co, Ni and Zn. Modelling solution iron as iron oxide particles decreased the log K(D,pred) values for Pb. The Cu predictions were very sensitive to the metal binding strength of dissolved and particulate organic matter. Overall, the model shows promise for the prediction of solid-solution metal partitioning in aquatic systems. The modelling exercise has identified the following uncertainties: the extent of Ca and Mg competition for binding sites, the chemical nature of the measured particulate metal, the efficacy of the solid-solution separation method and the strength of copper binding to organic matter.

Journal Article↗

Analysis of zinc fingers optimized via phage display: evaluating the utility of a recognition code.

Cys2His2 zinc finger proteins are composed of modular DNA-binding domains and provide an excellent framework for the design and selection of proteins with novel site specificity. Crystal structures of zinc finger-DNA complexes have shown that many Cys2His2 zinc fingers use a conserved docking arrangement that juxtaposes residues at key positions in the "recognition helix" with corresponding base positions in the three to four base-pair subsite. Several groups have proposed that specificity can be explained with a zinc finger-DNA recognition code that correlates specific amino acids at these key positions in the alpha-helix with specific bases in each position of the corresponding subsite. Here, we explore the utility of such a code through detailed studies of zinc finger variants selected via phage display. These proteins provide interesting systems for detailed analysis since they have affinities and specificities for their sites similar to those of naturally occurring DNA-binding proteins. Comparisons are facilitated by the fact that only key DNA-binding residues are varied in each finger while leaving all other regions of the structure unchanged. We study these proteins in detail by (1) selecting their optimal binding sites and comparing these binding sites with sites that might have been predicted from a code; (2) by examining the "evolutionary history" of these proteins during the phage display protocol to look for evidence of context-dependent effects; and (3) by reselecting finger 1 in the presence of the optimized finger 2/finger 3 domains to obtain further data on finger modularity. Our data for optimized fingers and binding sites demonstrate a clear correlation with contacts that would be predicted from a code. However, there are enough examples of context-dependent effects (not explained by any existing code) that selection is the most reliable method for maximizing the affinity and specificity of new zinc finger proteins.

Amino Acid Sequence↗

Identification of patients with Churg-Strauss syndrome (CSS) using automated data.

PURPOSE: Our aim was to identify individuals with Churg-Strauss syndrome (CSS) among asthma drug users, based on patterns of diagnostic and procedural codes (termed 'algorithms') contained in automated claims data. METHODS: A retrospective study was conducted among patients who had been dispensed asthma drugs at three HMOs. Individuals who received > or =3 dispensings of an asthma drug during any consecutive 12-month period beginning 1 January 1994 through 20 June 2000 were identified. Information on patient age, gender, enrollment status, asthma drugs dispensed, inpatient and outpatient diagnoses and procedures were obtained from the HMO automated databases. Twelve combinations of diagnostic and billing codes ('algorithms') were developed using the claims data to identify potential cases of CSS. Chart reviews blinded to drug exposure were performed using a standardized abstraction form. A rheumatologist reviewed abstracted information on all subjects, and those who met two or more American College of Rheumatology (ACR) criteria for CSS were further reviewed by two clinical experts. Cases were classified as unlikely, possible, or probable/definite CSS. Each clinical expert independently rated the cases; disagreements were resolved by consensus. RESULTS: A total of 185 604 patients who had been dispensed asthma drugs were identified. Three hundred fifty subjects were selected for chart review, and 15 were classified as having 'probable/definite' CSS. The algorithms that were most successful in identifying patients with CSS were as follows: (1) two or more codes for vasculitis (13 confirmed cases from 129 reviewed; positive predictive value 10%); (2) codes for both vasculitis and neurologic symptoms (6 confirmed cases from 15 reviewed; positive predictive value 40%) and (3) codes for both eosinophilia and vasculitis (4 confirmed cases from 5 reviewed; positive predictive value 80%). CONCLUSION: Automated claims data can be used to identify patients with CSS. This approach can facilitate better epidemiologic study of the risk factors for the condition.

Algorithms↗

Chromosomal level genome assembly of medicinal plant Chrysosplenium macrophyllum.

Chrysosplenium macrophyllum Oliv., a perennial herb native to China, is widely used in traditional medicine for its notable therapeutic properties. However, the absence of a reference genome has constrained its full potential for research and application. This study presents the first chromosome-level de novo genome assembly of C. macrophyllum, constructed by integrating long reads from Oxford Nanopore Technologies (ONT), short reads from BGI, and Hi-C data. The final assembly spans 2.55 Gb, with a scaffold N50 of 93.38 Mb, and 83.70% of the genome has been assigned to 22 chromosomes. The mapping rate of the BGI short reads to the genome is approximately 97.94%, and BUSCO analysis reveals that 97.94% of the predicted genes are complete. A total of 62,921 protein-coding genes were predicted, with functional annotations for 93.67% of them. This chromosome-level genome assembly represents an important resource for expanding our understanding of Chrysosplenium species and supports future genomic studies and applications.

Genome, Plant↗

Variable deletion of exon 9 coding sequences in cystic fibrosis transmembrane conductance regulator gene mRNA transcripts in normal bronchial epithelium.

The predicted protein domains coded by exons 9-12 and 19-23 of the 27 exon cystic fibrosis transmembrane conductance regulator (CFTR) gene contain two putative nucleotide-binding fold regions. Analysis of CFTR mRNA transcripts in freshly isolated bronchial epithelium from 12 normal adult individuals demonstrated that all had some CFTR mRNA transcripts with exon 9 completely deleted (exon 9- mRNA transcripts). In most (9 of 12), the exon 9- transcripts represented less than or equal to 25% of the total CFTR transcripts. However, in three individuals, the exon 9- transcripts were more abundant, comprising 39, 62 and 66% of all CFTR transcripts. Re-evaluation of the same individuals 2-4 months later showed the same proportions of exon 9- transcripts. Of the 24 CFTR alleles in the 12 individuals, the sequences of the exon-intron junctions relevant to exon 9 deletion (exon 8-intron 8, intron 8-exon 9, exon 9-intron 9, and intron 9-exon 10) were identical except for the intron 8-exon 9 region sequences. Several individuals had varying lengths of a TG repeat in the region between splice branch and splice acceptor consensus sites. Interestingly, one allele in each of the two individuals with 62 and 66% exon 9- transcripts had a TT deletion in the splice acceptor site for exon 9. These observations suggest either the unlikely possibility that sequences in exon 9 are not critical for the functioning of the CFTR or that only a minority of the CFTR mRNA transcripts need to contain exon 9 sequences to produce sufficient amounts of a normal CFTR to maintain a normal clinical phenotype.

Base Sequence↗

Correlation approach to identify coding regions in DNA sequences.

Recently, it was observed that noncoding regions of DNA sequences possess long-range power-law correlations, whereas coding regions typically display only short-range correlations. We develop an algorithm based on this finding that enables investigators to perform a statistical analysis on long DNA sequences to locate possible coding regions. The algorithm is particularly successful in predicting the location of lengthy coding regions. For example, for the complete genome of yeast chromosome III (315,344 nucleotides), at least 82% of the predictions correspond to putative coding regions; the algorithm correctly identified all coding regions larger than 3000 nucleotides, 92% of coding regions between 2000 and 3000 nucleotides long, and 79% of coding regions between 1000 and 2000 nucleotides. The predictive ability of this new algorithm supports the claim that there is a fundamental difference in the correlation property between coding and noncoding sequences. This algorithm, which is not species-dependent, can be implemented with other techniques for rapidly and accurately locating relatively long coding regions in genomic sequences.

Algorithms↗

Comparisons of defective HTLV-I proviruses predict the mode of origin and coding potential of internally deleted genomes.

Cell lines infected with a variety of HTLV-I isolates were examined for the presence of defective proviruses that contain deletions spanning the gag, pol, and env genes. Internally deleted proviruses were identified by Southern blotting and by PCR amplification with 5' and 3' primers complementary to gag and tax sequences, respectively. PCR products representing eight defective proviruses from seven different cell lines were subsequently cloned and sequenced. The objectives of this study were twofold: first, we sought to determine whether nucleotide sequences surrounding sites of deletion shared common features that might reveal the mechanisms by which the defective genomes originated. Second, we asked whether deleted proviruses encode Gag fusion proteins with related C-terminal residues derived from open reading frames in the pX region. While most of the defective proviruses had incurred a single, large deletion, two of them displayed a more complex pattern of multiple rearrangements. Alignments of bases flanking the 5' and 3' deletion endpoints within each provirus showed tracts of sequence identity consistent with a mechanism involving aberrant intramolecular strand-transfer events during replication. We suggest that the amount or activity of HTLV-I polymerase in virions may contribute both to the poor infectivity of the virus and to the high deletion frequency. Two of the eight proviruses that were examined encoded a gag gene joined to an extended open reading frame; the other six had very short open reading frames (one to six amino acids) derived from pX or env regions joined to gag that showed no apparent amino acid sequence similarity.

Amino Acid Sequence↗

Control of remembered reaching sequences in monkey. I. Activity during movement in motor and premotor cortex.

Motor and premotor cortex firing patterns from 307 single neurons were recorded while monkeys made rapid sequences of three reaching movements to remembered target buttons arrayed in two-dimensional space. A primary goal was to study and compare directionally tuned responses for each of three movement periods during 12 movement sequences that uniformly sampled the directional space in front of the monkey. The majority of neurons showed maximal responses during movements in a preferred direction with smaller increases during movements close to the preferred direction. These responses showed a statistically significant regression fit to a cosine function for 72% of the neurons examined. Comparisons among tuning directions computed separately for the first, second, and third movement periods suggested the near constancy of preferred direction across a rapidly executed series of movements even though these movements began at different starting points in space. Although directionally tuned neurons were only broadly tuned for a specific direction of movement, the neuronal ensemble carried accurate directional information. A population vector computed by summing vector contributions from the entire population of tuned neurons predicted movement direction with a mean accuracy of 20 degrees. This population code made consistent predictions for each of the 36 movements that were studied using a single set of population parameters. Most of the remaining neurons (24%) that were not tuned during movement did show significant changes in activity during other aspects of task performance. Some nontuned neurons had nondirectional increases that were sustained during movement, while others showed identical phasic bursts during the three movement periods. These nontuned neurons may control stabilizations of the shoulder, trunk, and forearm during movement, or forearm movements during button pushing.

Animals↗