Search PubMedSearch

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Replacing tracheoesophageal voicing sources using LPC synthesis.

The feasibility of using the linear predictive coding (LPC) technique to replace the voicing sources of tracheoesophageal speech was explored. Four vowels, [i], [a], [e], [u], and one diphthong [ou], produced by two male and two female tracheoesophageal speakers were analyzed by the LPC autocorrelation method. Normalized prediction error functions were used to choose the algorithm and the control parameters of the analysis. Poles of the vocal tract transfer function were selected from frames whose normalized prediction errors were close to minimum with criteria derived from transfer functions of normally produced vowels. Vowels were synthesized with the reconstructed transfer function and a synthesized excitation input. Results of an identification task indicated that the synthesized vowels were highly intelligible, and the gender of the speaker was better identified from the synthesized vowels than from the original tracheoesophageal vowels.

Female

BIGPROBE: a computer program that predicts the sequence of long oligonucleotide probes with high reliability.

We have written a computer program, BIGPROBE, which facilitates the design of long nucleic acid probes from the partial or complete amino acid sequence of a protein. BIGPROBE relies upon information on codon usage, intercodon dinucleotide frequency, and potential probe self-complementarity. We have examined the accuracy with which the program predicts coding sequences using sample human and rat genes and probe lengths of 30-60 nucleotides. Rat probe sequences selected by BIGPROBE using either codon usage or dinucleotide frequency data alone averaged 86-92% homology with the known exons of the corresponding gene sequences. Predictive accuracy with rat gene probes could be improved to 89-94%, depending upon probe length, by applying codon usage and dinucleotide frequency data in combination. Similar accuracy was achieved for human genes.

Amino Acid Sequence

Acoustic correlates of vocal quality.

We have investigated the relationship between various voice qualities and several acoustic measures made from the vowel /i/ phonated by subjects with normal voices and patients with vocal disorders. Among the patients (pathological voices), five qualities were investigated: overall severity, hoarseness, breathiness, roughness, and vocal fry. Six acoustic measures were examined. With one exception, all measures were extracted from the residue signal obtained by inverse filtering the speech signal using the linear predictive coding (LPC) technique. A formal listening test was implemented to rate each pathological voice for each vocal quality. A formal listening test also rated overall excellence of the normal voices. A scale of 1-7 was used. Multiple linear regression analysis between the results of the listening test and the various acoustic measures was used with the prediction sums of squares (PRESS) as the selection criteria. Useful prediction equations of order two or less were obtained relating certain acoustic measures and the ratings of pathological voices for each of the five qualities. The two most useful parameters for predicting vocal quality were the Pitch Amplitude (PA) and the Harmonics-to-Noise Ratio (HNR). No acoustic measure could rank the normal voices.

Adult

The qa repressor gene of Neurospora crassa: wild-type and mutant nucleotide sequences.

The qa-1S gene, one of two regulatory genes in the qa gene cluster of Neurospora crassa, encodes the qa repressor. The qa-1S gene together with the qa-1F gene, which encodes the qa activator protein, control the expression of all seven qa genes, including those encoding the inducible enzymes responsible for the utilization of quinic acid as a carbon source. The nucleotide sequence of the qa-1S gene and its flanking regions has been determined. The deduced coding sequence for the qa-1S protein encodes 918 amino acids with a calculated molecular weight of 100,650 and is interrupted by a single 66-base-pair intervening sequence. Both constitutive and noninducible mutants occur in the qa-1S gene and two different mutations of each type have been cloned and sequenced. All four mutations occur within the predicted coding region of the qa-1S gene. This result strongly supports the hypothesis that the qa-1S gene encodes a repressor. All four mutations are located within codons for the last 300 amino acids of the qa-1S protein. The mutations in three of the mutants involve amino acid substitutions, while the fourth mutant, which has a constitutive phenotype, contains a frameshift mutation. The two constitutive mutations occur in the most distal region of the gene, possibly implicating the COOH-terminal region of the qa repressor in binding to its target. The two noninducible mutations occur in a region proximal to the constitutive mutations, possibly implicating this region of the qa repressor in binding the inducer.

Base Sequence

Cloning and sequence comparison of the mouse, human, and chicken engrailed genes reveal potential functional domains and regulatory regions.

We have isolated and characterized genomic DNA clones for the human and chicken homologues of the mouse En-1 and En-2 genes and determined the genomic structure and predicted protein sequences of both En genes in all three species. Comparison of these vertebrate En sequences with the Xenopus En-2 [Hemmati-Brivanlou et al., 1991) and invertebrate engrailed-like genes showed that the two previously identified highly conserved regions within the En protein ]reviewed in Joyner and Hanks, 1991] can be divided into five distinct subregions, designated EH1 to EH5. Sequences 5' and 3' to the predicted coding regions of the vertebrate En genes were also analyzed in an attempt to identify cis-acting DNA sequences important for the regulation of En gene expression. Considerable sequence similarity was found between the mouse and human homologues both within the putative 5' and 3' untranslated as well as 5' flanking regions. Between the mouse and Xenopus En-2 genes, shorter stretches of sequence similarity were found within the 3' untranslated region. The 5' untranslated regions of the mouse, chicken and Xenopus En-2 genes, however, showed no similarly conserved stretches. In a preliminary analysis of the expression pattern of the human En genes, En-2 protein and RNA were detected in the embryonic and adult cerebellum respectively and not in other tissues tested. These patterns are analogous to those seen in other vertebrates. Taken together these results further strengthen the suggestion that En gene function and regulation has been conserved throughout vertebrate evolution and, along with the five highly conserved regions within the En protein, raise an interesting question about the presence of conserved genetic pathways.

Amino Acid Sequence

Expression of the FGF-related proto-oncogene int-2 during gastrulation and neurulation in the mouse.

The proto-oncogene int-2 has been implicated in the formation of mouse mammary-tumour-virus-induced mammary tumours. Analysis of the predicted coding sequence indicates that int-2 is a member of the fibroblast growth factor family. Previous studies using Northern blot analysis suggested that normal expression of int-2 may be confined to extra-embryonic endoderm lineages of embryonic stages of mouse development. We have used in situ hybridization and Northern blot analysis to examine directly int-2 expression in embryo stem cells and in the developing embryo from early gastrulation to midsomite stages. Complex patterns of accumulation of int-2 RNA were observed in embryonic and extra-embryonic tissues. The data suggest multiple roles for int-2 in development which may include migration of early mesoderm cells and induction of the otocyst.

Animals

Compact digital storage of ECG's.

The technique of predictive coding is applied to the problem of reversible compression of digitized electrocardiograms. Integer-based predictors and MMSE predictors are studied as regards performance at varying sampling rates and digital resolutions for both long-term ECGs and ECGs recorded at rest. It is concluded that MMSE predictors are to be preferred only in the case when the ECG is oversampled (i.e., the sampling rate is much higher than twice the cut-off frequency of the presampling filter). In other cases the integer predictor which yields the so-called 2nd differences is superior. The problem of encoding the resulting residuals with a variable-length code is studied for long-term ECGs digitized at 100 Hz and using 8 bits digital resolution. The code has a simple struture leading to speed of execution while the efficiency loss is small.

Computers

The nucleoprotein gene of Ebola virus: cloning, sequencing, and in vitro expression.

Genomic and messenger RNAs of a Zaire strain of Ebola virus were cloned, and inserts specific for the nucleoprotein gene were isolated and sequenced. The nucleoprotein gene is located proximal to the 3' end of the genome and is preceeded by a putative leader sequence. The gene begins with the transcriptional start site sequence 3'-UACUCCUUCUAAUU..., and ends with the polyadenylation site sequence 3'-... UAAUUCUUUUUU. The predicted coding region is 2217 bases in length and encodes a protein that contains 739 amino acids, with a calculated molecular weight of 83.3 kDa. The protein has an approximate net charge of -30 and can be divided into a hydrophobic N-terminal half and a hydrophilic and highly acidic C-terminal half. An in vitro transcript, generated from plasmid DNA containing the entire coding region, directs the synthesis of authentic nucleoprotein in a rabbit reticulocyte lysate system. The genomic organization and transcriptional signals of Ebola are similar to those of other nonsegmented, negative-strand RNA viruses, but nucleic acid or amino acid sequence comparisons indicate a lack of similarity.

Amino Acid Sequence

An acoustic study of vowel production in aphasia.

A group of five anterior and seven posterior aphasic patients were recorded for their vowel productions of the nine nondipthong vowels of American English and compared to a group of seven normal speakers. All phonemic substitutions were eliminated from the data base. A Linear Predictive Coding (LPC) computer program was used to extract the first and the second formant frequencies at the midpoint of the vowel for each of the remaining repetitions of the nine vowels. The vowel duration and the fundamental frequency of phonation were also measured. Although there were no significant differences in the formant frequency means across groups, there were significantly larger standard deviations for the aphasic groups compared to normals. Anterior aphasics were not significantly different from posterior aphasics with respect to this greater formant variability. There was a main effect for vowel duration means, but no individual group was significantly different from the other. Standard deviations of duration were significantly greater for the anterior aphasics compared to normal speakers, but not significantly different from posterior aphasics. Posterior aphasics did not have significantly greater standard deviations of duration than did normal subjects. Greater acoustic variability was considered to evidence a phonetic production deficit on the part of both groups of aphasic speakers, in the context of fairly well-preserved phonemic organization for vowels.

Adult

Anticipatory coarticulation in aphasia: acoustic and perceptual data.

Two experiments investigated speech motor planning in aphasia by contrasting the degree of labial and lingual anticipatory coarticulation evident in normal subjects' speech with that found in the speech of aphasic subjects. In the first experiment, Linear Predictive Coding (LPC) analyses were conducted for the initial consonants of CV [si su ti tu ki ku] and CCV [sti stu ski sku] productions by 6 normal and 10 aphasic (5 anterior, 5 posterior) subjects. For normal subjects' productions, reliable coarticulatory shift was found for almost all measurements, indicating that acoustic correlates for anticipatory coarticulation obtain for [s], [t], and [k] in a prevocalic environment, as well as when [s] is the initial consonant of a CCV syllable. The data for the aphasic subjects were statistically indistinguishable from those of the normal subject group, and there were no differences noted as a function of aphasia type. In the second experiment, a subset of the consonantal stimuli produced by the normal and aphasic subjects was presented to a group of 10 naive listeners for a vowel identification task. Listeners were able to identify the productions of all subjects at a level well above chance. In addition, small but statistically significant Group differences were observed, with the [sV], [skV], and [tV] productions by anterior aphasics showing significantly lower perceptual scores than those of normal subjects.

Aged

Spectral analysis of heart rate variability following human heart transplantation: evidence for functional reinnervation.

To determine the status of innervation in long-term human donor allografts, the power spectrum of heart rate variability was analysed in 9 post-transplant patients and 7 healthy control subjects. The mean post-transplant follow-up was 17.8 months (range: 2-37 months). Continuous ECG signals were recorded throughout a 15-min rest period. An R-R interval tachogram was generated and an autoregressive model using linear predictive coding, was applied to the heart rate variability data. In 8 transplant patients the frequency oscillations were irregular, broad based and widely dispersed from 0 to 1 Hz. The patterns resembled white noise and were consistent with dissociation of the donor allograft from the recipient's central nervous system. In contrast, one patient displayed a heart rate variability spectrum indistinguishable from that of control subjects. This pattern contained two distinct spectral bands; one corresponding to the patient's respiratory rate at 0.2 Hz and a low frequency Mayer wave at 0.1 Hz. Atropine abolished the respiratory (vagal) peak. Except for this patient's post-transplant time (33 months compared to the group mean of 17.6 months), there were no clinical characteristics which distinguished this patient from the others. While the mean heart rate for the remaining 8 allografts was significantly higher than controls (95.3 vs 64.5 bt/min; P less than 0.001) the standard deviation of heart rate variability for the 8 patients was significantly narrower than controls (0.7 vs 4.86; P less than 0.01). The variance of heart rate for the patient with the normal power spectrum was fourfold greater than the mean SD of the other transplant patients.(ABSTRACT TRUNCATED AT 250 WORDS)

Adult

P-glycoprotein genes in the winter flounder, Pleuronectes americanus: isolation of two types of genomic clones carrying 3' terminal exons.

In mammals, P-glycoprotein (P-gp) is encoded by two or more highly conserved genes that differ in their abilities to transport drugs. One isoform class (class I) is consistently associated with the multidrug resistance phenotype, while the other (class III) is not. This study was designed to enumerate the P-gp genes in fish and determine how they are related to the two functional classes already defined in mammals. Southern blot analysis using a conserved single exon from the 3' terminal region of hamster P-gp cDNA (pEX1-172) as a probe indicated that there were two P-gp genes in right-eye flounders. Subsequently, two sets of clones were isolated from a winter flounder genomic library that correspond to the 3' ends of the two flounder P-gp genes. Sequence analysis was done on two key areas: the 3' ATP binding site and the 3' terminal exon, both of which were found to be homologous with their mammalian counterparts. Despite high levels of sequence identity in the predicted coding regions of the gene fragments it has not been possible to use these sequences to relate the homologs to particular mammalian classes of P-gp genes, perhaps because of gene conversion between mammalian P-gp genes. These cloned sequences are the first set of P-gp genes reported in lower vertebrates and will be useful for delineating the expression of P-gp genes in fish and understanding the role of P-gp in fish physiology.

ATP Binding Cassette Transporter, Subfamily B, Mem

Methods for the identification of evoked response components in the frequency and combined time/frequency domains.

Two prominent frequency components designated f1 and f2 have been identified in the visual evoked response to the transient presentation of sinusoidal luminance gratings in the range of 0.5-8 c/deg. The components occur at temporal frequencies below the alpha band, with the f1 frequency being roughly half that of the f2 frequency. The f1 component is largest at low spatial frequencies with f2 becoming progressively dominant as spatial frequency is increased. The frequency and amplitude of f1 and f2 change substantially over the time course of the response. This has been studied by calculating the temporal frequency spectrum of the transient evoked potential over successive short-time epochs running through the response. Using this technique, the response is shown to consist of narrow-band frequency peaks or 'formants' emerging at different times after stimulus onset. These formants occur at frequencies other than those of the spontaneous EEG and undergo changes in frequency and amplitude over the time course of the response. Two spectrum analysis techniques were employed: the Discrete Fourier Transform and Linear Predictive Coding. Frequency components were successfully identified in single-trial responses using the LPC technique.

Electroencephalography

The Personal Acoustics Lab (PAL): a microcomputer-based system for digital signal acquisition, analysis, and synthesis.

A new, integrated digital signal processing (DSP) system, the Personal Acoustics Lab (PAL), is described. This microcomputer-based system is suitable for analogue signal digitization at rates from several samples per hour to 150,000 samples per second in 12- or 16-bit words. Data may be acquired on one to sixteen single-ended A/D, or one to eight double-ended A/D channels in bipolar or unipolar modes. Digitized data may be reconverted to analogue signals using one or two D/A channels. An external clock and trigger and two bidirectional digital ports are provided. Integrated PAL-ILS software commands perform all necessary DSP functions, including: data editing, time- and frequency-domain graphical display, plotting, filtering, Fourier and Hilbert transforms, linear predictive coding, auto- and cross-correlation, and summary statistics. The system is suitable for biological and engineering DSP applications. Output from selected PAL-ILS software commands is illustrated using a bioacoustical example.

Animals

Characterization of a G-protein alpha-subunit gene from the nematode Caenorhabditis elegans.

A gene encoding the alpha-subunit of a guanine nucleotide binding regulatory protein (G-protein) was isolated from a library of genomic Caenorhabditis elegans DNA. The predicted coding region is colinear to related genes from mammals and the 356 amino acid residues show 63% sequence identity to e.g. rat Gi alpha 2. Three of the eight introns within the coding sequence are at exactly the same positions as those in a Drosophila G-protein alpha-subunit gene, and two of these are also conserved in the mammalian homologues. The nematode gene does not encode the cysteine residue that forms the substrate site for pertussis toxin-catalyzed ADP-ribosylation in several G-proteins. In spite of the similarity to mammalian G-protein alpha-subunit genes the gene can not unambiguously be categorized in one of the classes of G-proteins recognized in mammals (G alpha i, o, z, etc.). The position of the gene on the physical map of the animal was determined (chromosome V). The cloning and sequencing of this gene can be the starting point of reverse genetics experiments aimed at the isolation of animals mutated in a G-protein alpha-subunit gene.

Amino Acid Sequence

MetagenomicKG: a knowledge graph for metagenomic applications.

MOTIVATION: The sheer volume and variety of genomic content within microbial communities makes metagenomics a field rich in biomedical knowledge. To traverse these complex communities and their vast unknowns, metagenomic studies often depend on distinct reference databases, such as the Genome Taxonomy Database (GTDB), the Kyoto Encyclopedia of Genes and Genomes (KEGG), and the Bacterial and Viral Bioinformatics Resource Center (BV-BRC), for various analytical purposes. These databases are crucial for the genetic and functional annotation of microbial communities. Nevertheless, the inconsistent nomenclature or identifiers of these databases present challenges for effective integration, representation, and utilization. Knowledge graphs (KGs) offer an appropriate solution by organizing biological entities from different databases to standardized identifiers, allowing their interrelations to be captured into a cohesive network regardless of the naming conventions used in each source. The graph structure not only facilitates the unveiling of hidden patterns but also enriches our biological understanding with deeper insights. Despite KGs having shown potential in various biomedical fields, their application in metagenomics remains underexplored. RESULTS: We present MetagenomicKG, a novel knowledge graph specifically tailored for metagenomic analysis. MetagenomicKG integrates taxonomic, functional, and pathogenesis-related information on the human microbiome sourced from various databases, and further connects these with existing biomedical KGs to expand the biological network. Through various case studies involving the human microbiome, we demonstrate its utility in enabling hypothesis generation regarding the relationships between microbes and diseases, generating sample-specific graph embeddings, and providing robust pathogen prediction. CODE AVAILABILITY: The source code and technical details for constructing the MetagenomicKG and reproducing all analyses are available on GitHub at https://github.com/KoslickiLab/MetagenomicKG. The data used in this manuscript, including the pre-built files and use case input data, are archived on Zenodo with DOI: 10.5281/zenodo.17546861.

Metagenomics

Sequence of a pseudogene in the legumin gene family of pea (Pisum sativum L.).

A second legumin gene, denoted psi Leg D, has been located on the pea genomic clone lambda Leg 1, approx. 1.3 Kbases 3' of Leg A, in the same orientation. The complete sequence of psi Leg D shows that it is a pseudogene, having two stop codons near the 5' end of its predicted coding sequence, as well as deletions and frame shift errors when compared to Leg A. No transcripts from this gene could be detected in developing pea seeds. Leg A and psi Leg D are homologous over their coding sequences, and partially homologous in the intron sequences and the immediate 5' flanking sequences. Other flanking sequences of the two genes show no significant homology, apart from the presence of polyadenylation signals 3' to both coding sequences. The introns in the two genes occur in corresponding positions in the sequences, but a deletion in psi Leg D affects the 3' boundary of IVS-2. Hybridisation of psi Leg D to pea genomic DNA suggests that it does not represent a hitherto undetected sub-family of legumin genes.

Amino Acid Sequence

Retrotransposon-like nature of Tp1 elements: implications for the organisation of highly repetitive, hypermethylated DNA in the genome of Physarum polycephalum.

The repetitive fraction of the genome of the eukaryotic slime mould Physarum polycephalum is dominated by the Tp1 family of highly repetitive retrotransposon-like sequences. Tp1 elements consist of two terminal direct repeats of 277bp which flank an internal domain of 8.3kb. They are the major sequence component in the hypermethylated (M+) fraction of the genome where they have been found exclusively in scrambled clusters of up to 50kb long. Scrambling is thought to have arisen by insertion of Tp1 into further copies of the same sequence. In the present study, sequence analysis of cloned Tp1 elements has revealed striking homologies of the predicted amino acid sequence to several highly conserved domains characteristic of retrotransposons. The relative order of the predicted coding regions indicates that Tp1 elements are more closely related to copia and Ty than to retroviruses. Self-integration and methylation of Tp1 elements may function to limit transposition frequency. Such mechanisms provide a possible explanation for the origin and organisation of M + DNA in the Physarum genome.

Amino Acid Sequence