Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sound Spectrography”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 937 records · Page 52Linked to original sources

The relation between vowel recognition and measures of frequency resolution.

The purpose of this study was to employ measures of frequency resolution obtained from individual subjects to predict each subject's vowel recognition performance. Input filter patterns at six test frequencies were obtained from normal-hearing and hearing-impaired subjects. These patterns were used to correlate frequency resolution with vowel recognition in those same subjects. Vowels were presented at levels at which the entire spectrum was fully audible to each subject. Using each subject's measured filter characteristics (and interpolated values for intermediate frequencies), an "internal spectrum" of each vowel was calculated by determining the outputs of all filter channels for the vowel as the input signal. It was speculated that the more similar two internal spectra for a subject were, the more often they would be confused in the vowel recognition task. This expectation received some support when the measure of similarity was a point-by-point Euclidean distance between the two internal spectra. Stronger support was obtained when the measure of similarity was based upon Klatt's (1982) "weighted slope metric" that emphasizes similarities of spectral peak locations. The present study demonstrates a relation between impairments of frequency resolution and vowel recognition. The described filter-bank model of vowel recognition suggests that measures of frequency resolution along with the acoustic spectra of vowel stimuli may be useful in predicting the recognition of vowels by individuals.

Adult↗

Comparison of two single-channel vibrotactile aids for the hearing-impaired.

Two commercially available single-channel vibrotactile aids, designed to transmit information about acoustic stimuli to persons who cannot perceive such stimuli through conventional amplification, were compared in a number of tasks with the same subjects. Both devices employed a vibratory transducer worn on the wrist. One device represented characteristics of the envelope of the waveform by using it to modulate the amplitude of a 250-Hz carrier vibration (an amplitude-modulated, or AM, signal). The other device presented and amplitude-modulated a broad-band signal whose spectral characteristics preserved information about the signal. Subjects performed several tasks. On some tasks (sound detection, environmental sound identification, syllable rhythm and stress categorization) information about the envelope of the stimulus was expected to be sufficient for good performance. On others (speech sound recognition) additional information about the spectral fine structure of the signal spectrum was anticipated to be required for good performance. Results indicated that the subjects performed comparably with both devices on all tasks, suggesting that they did not make use of the spectral information available in the more complex signal.

Adolescent↗

Spectral correlates of glottal voice source waveform characteristics.

The relationships between the waveform and the spectrum of the pulsating transglottal airflow during vowel phonation are analyzed in singers and nonsingers. The waveform, called the flow glottogram, is analyzed by means of inverse filtering, and the spectrum is determined either directly, by submitting the flow glottogram to spectrum analysis, or indirectly, by measuring spectral changes accompanying phonatory changes under conditions of constant vowel articulation. The peak-to-peak amplitude of the flow glottogram pulses shows a strong relationship with the amplitude of the source spectrum fundamental and varies considerably during phonation, presumably depending on the degree of glottal ab/adduction. The negative peak amplitude of the differentiated flow glottogram shows a high correlation with the sound pressure level of the vowel.

Adult↗

Statistical differentiation of tracheoesophageal speech produced under four prosthetic/occlusion speaking conditions.

Twelve male and 12 female total laryngectomy patients who received the tracheoesophageal puncture (TEP) as a means of vocal rehabilitation served as subjects for this investigation. Recordings were made of these subjects' speech produced with four prosthetic/occlusion conditions: (1) duckbill prosthesis with tracheostoma valve; (2) duckbill prosthesis with digital occlusion of the tracheostoma; (3) low pressure prosthesis with tracheostoma valve; and (4) low pressure prosthesis with digital occlusion. Speech tasks consisted of three trials of maximum phonation time on /a/ and reading of a 98-word standard passage. Acoustic analysis of the recorded speech samples included a total of 34 frequency, intensity, temporal, and noise measures. Eight acoustic measures (words per minute, harmonics-to-noise ratio, percent jitter, intensity range during vowel phonation, percent periodic phonation, mean intensity during reading, directional jitter, and directional jitter, and directional shimmer) were chosen as dependent variables for a repeated measures MANOVA. The overall repeated measures MANOVA, a set of complex contrasts, and paired t tests revealed that TEP speech produced with the low pressure prosthesis was significantly different from that produced with the duckbill prosthesis on a weighted linear combination of the eight acoustic variables. Tracheoesophageal voice produced with a low pressure prosthesis had greater amounts of periodic phonation than tracheoesophageal voice produced with a duckbill prosthesis. The use of a tracheostoma valve did not have a significant impact on the subset of acoustic measures used in the repeated measures MANOVA.

Female↗

Some effects of variations in response time latency on speech rate, interruptions, and fluency in children's speech.

The present study was designed to examine adult-child interactions during conversation with respect to the effects of adult paralinguistic speech variations on the speech production of children. Four 4-year-old children served as subjects. A single-subject A-B-A design with counterbalancing and replication was implemented. Each subject participated in three 15-min conversations with an experimenter. The independent variable was the interspeaker pause time--the response time latency (RTL). During the 15-min conversations, the experimenter used either a 1-s or a 3-s RTL when responding to the child. RTL was measured for each subject in each condition. Data analysis revealed that each child's RTL was significantly longer when the experimenter's RTL was 3 s than when it was 1 s, and all differences between all conditions reached significance for these subjects. Other dependent variables included speech rate, the frequency of disfluencies, and the frequency of interruptions produced by the subjects within each condition. All 4 subjects varied the frequency of disfluencies and interruptions. However, each child varied rate and disfluencies in a highly individualistic manner.

Child, Preschool↗

Acoustic measurements of objective tinnitus.

Ear canal sound pressure levels were measured from a 38-year-old woman who had experienced objective tinnitus in her right ear for approximately 2 years. The tinnitus sounded like a series of "sighs" that were synchronous with her pulse rate. Because the level of the tinnitus fluctuated in a pulsing manner, it appeared to be of vascular origin. Psychoacoustically, the tinnitus behaved like a low-pass masker (cutoff frequency = 1.5 kHz) of about 40 dB SPL. This masking effect was manifested as a low-frequency hearing loss in the subject's right ear. A miniature microphone system was used to monitor the tinnitus before, during, and after a jugular-vein ligation. Because the cause of the tinnitus was only generally known, acoustically monitoring the sound as the jugular vein and/or its tributaries were systematically clamped and then released enabled the site of generation to be known exactly. By monitoring the tinnitus during surgery, the effectiveness of the corrective procedure could be immediately evaluated. Hearing sensitivity in the affected ear returned to normal limits following the elimination of the tinnitus. One year after the surgery, the tinnitus was barely audible to the woman, but only when she positioned her head a specific way. The level of the tinnitus measured in this head-turned condition was markedly lower than the level obtained preoperatively.

Adult↗

Acoustic correlates of pathologic voice types.

Listeners classified 49 samples of vowels /a/ and /i/ on the basis of four voice types: hoarse, breathy, strained, and normal. The vowels were analyzed acoustically for mean harmonic/noise differences in four spectral regions, average fundamental frequency, natural logarithm of fundamental frequency, and jitter. Discriminant analysis showed that classifications of voice type were made with 80% accuracy using three acoustic parameters: (a) mean harmonic/noise difference factor (1-3.5 kHz), (b) natural log of fundamental frequency, and (c) vowel type. The significance of these particular acoustic parameters for the perception and classification of voice types is discussed.

Discriminant Analysis↗

Laryngeal perturbation analysis: minimum length of analysis window.

Laryngeal perturbation measures have been applied to the analysis of cycle-to-cycle changes in periodicity and amplitude of the acoustic voice signal for more than 25 years. Although such measures enjoy widespread clinical application, there is little agreement about basic methodology, including the length of signal to be analyzed. The purpose of this study was to examine changes in laryngeal perturbation measures as a function of length of signal analyzed in 18 subjects who complained of symptoms of possible laryngeal dysfunction. The results showed that as many as 190 cycles may be necessary before jitter asymptotes and as many as 130 cycles may be necessary before shimmer asymptotes. Pathological voices may require a longer analysis window for perturbation analysis than do nonpathological voices.

Adult↗

Pitch effects on vowel roughness and spectral noise for subjects in four musical voice classifications.

This study was designed to investigate the effects of vocal fo on vowel spectral noise level (SNL) and perceived vowel roughness for subjects in high- and low-pitch voice categories. The subjects were 40 adult singers (10 each sopranos, altos, tenors, and basses). Each produced the vowel /a/ in isolation at a comfortable speaking pitch, and at each of seven assigned pitches spaced at whole-tone intervals over a musical octave within his or her singing pitch range. The eight /a/ productions were repeated by each subject on a second test day. The SNL differences between repeated test samples (different days) were not statistically significant for any subject group. For the vowel samples produced at a comfortable pitch, a relatively large SNL was associated with samples phonated by the subjects of each sex who manifested the relatively low singing pitch range. Regarding the vowel samples produced at the assigned-pitch levels, it was found that both vowel SNL and perceived vowel roughness decreased as test-pitch level was raised over a range of one octave. The relationship between vocal pitch and either vowel roughness or SNL approached linearity for each of the four subject groups.

Acoustics↗

Assessment of the dynamics of vocal fold contact from the electroglottogram: data from normal male subjects.

Electroglottographic (EGG) and acoustic records from 10 normal men prolonging the vowel /a/ at 60-68 dB, 70-78 dB, and 80-88 dB SPL were obtained. "Contact quotient" (EGG duty cycle) was shown to vary directly with vocal SPL. The mean contact quotient was 0.57 (SD = 0.07) and varied on the order of 1% over the course of a given phonation. "Contact index", a metric of EGG symmetry, also tended to vary with SPL. Consistent with previous qualitative descriptions of EGG morphology in modal register voice, the contact index averaged -0.52 (SD = 0.08), indicating that the EGG "closing phase" represents about 24% of the entire "contact phase". Contact index was more variable than contact quotient on consecutive EGG waves, varying by about 10% during phonation. Subjects were also instructed to produce a slow crescendo. Sound pressure and EGG data indicated that both the slope of increasing EGG contact and EGG duty cycle were significantly related to the amplitude of the acoustic signal. These results suggest that quantitative electroglottography may provide powerful insights into the control and regulation of normal phonation and into the detection and characterization of pathology.

Adult↗

Low-frequency energy deficit in electrolaryngeal speech.

The present exploratory project was undertaken (a) to determine the relative strength of low-frequency energy in the output of one widely used electronic artificial larynx (Servox) and (b) to assess the relative strength of low-frequency energy in vowels produced by users of this type of artificial larynx. We hypothesized that the outputs of electronic artificial larynges and the vowels produced by laryngectomized users of these devices would be characterized by significant deficits in low-frequency energy level. Five users of electronic larynges and 10 normal speakers (5 female and 5 male) provided the speech samples. Results of spectral analyses indicated that there was a significant deficit in low-frequency energy both in the acoustic signals generated by a Servox electronic larynx and in vowels produced by laryngectomized users of this type of electronic larynx. Based on these findings, a second order filter was designed and implemented digitally to compensate for the observed deficit in low-frequency energy. A perceptual experiment was completed to evaluate the effect of low-frequency enhancement on perceived speech quality. Ninety-eight percent (+/- 2%) of the responses of listeners indicated that low-frequency enhanced speech samples had better vocal quality or were more pleasant to listen to than the original speech samples. We conclude that consideration for enhancing low-frequency characteristics is warranted in the design of improved prosthetic devices for alaryngeal speakers.

Aged↗

Anticipatory coarticulation in the speech of profoundly hearing-impaired and normally hearing children.

The present study investigated the extent of anticipatory coarticulation in the speech of five 7-year-old and four 10-year-old children with profound prelingual hearing impairment as compared to normally hearing age-matched control subjects. Ten tokens each of the CV syllables [integral of i, integral of u, ti, tu, ki, ku] were elicited from each of the children. Both temporal and spectral (centroid and F2 frequency) analyses were conducted to explore the influence of the following vocalic environment on the initial consonants. The data indicated that the hearing-impaired children displayed evidence of coarticulation on most measures, but they did so to a lesser degree when compared to the normally hearing children. The results are discussed in relation to theories of speech production in the hearing impaired, and their implications for the development of coarticulation are considered.

Child↗

Perseveratory coarticulation in the speech of profoundly hearing-impaired and normally hearing children.

This study investigated the extent of perseveratory coarticulation in the VC syllables [i integral of, u integral of, it, ut, ik, uk] as produced by 7- and 10-year-old normally hearing and profoundly hearing-impaired children. Measures of both temporal and spectral (centroid and F2 frequency) parameters were computed. The data revealed that the hearing-impaired speakers exhibited measurable but smaller effects of perseveratory coarticulation relative to the normally hearing speakers. These results are compared with studies of anticipatory coarticulation and are discussed in relation to the claim that perseveratory coarticulation is largely a result of the inertial properties of the speech production mechanism.

Child↗

Enhancement of word-recognition performance with a filtering technique.

The NU No. 6 materials spoken by a female speaker were passed through a notch filter centered at 247 Hz with a 34-dB depth. The filtering reduced the amplitude range within the spectrum of the materials by 10 dB that was reflected as a 7.5-vu reduction measured on a true vu meter. Thus, the notch filtering in effect changed the level calibration of the materials. Psychometric functions of the NU No. 6 materials filtered and unfiltered in 60-dB SPL broadband noise were obtained from 12 listeners with normal hearing. Although the slopes of the functions for the two conditions were the same, the functions were displaced by an average of 5.8 dB with the function for the filtered materials located at the lower sound-pressure levels.

Analog-Digital Conversion↗

Speech perception in adult subjects with familial dyslexia.

Speech perception was investigated in a carefully selected group of adult subjects with familial dyslexia. Perception of three synthetic speech continua was studied: /a/-/e/, in which steady-state spectral cues distinguished the vowel stimuli; /ba/-/da/, in which rapidly changing spectral cues were varied; and /sta/-/sa/, in which a temporal cue, silence duration, was systematically varied. These three continua, which differed with respect to the nature of the acoustic cues discriminating between pairs, were used to assess subjects' abilities to use steady state, dynamic, and temporal cues. Dyslexic and normal readers participated in one identification and two discrimination tasks for each continuum. Results suggest that dyslexic readers required greater silence duration than normal readers to shift their perception from /sa/ to /sta/. In addition, although the dyslexic subjects were able to label and discriminate the synthetic speech continua, they did not necessarily use the acoustic cues in the same manner as normal readers, and their overall performance was generally less accurate.

Dyslexia↗

Dysphonia detected by pattern recognition of spectral composition.

The vowel [a:] in a test word, judged normal or dysphonic, was examined with the Self-Organizing Map; the artificial neural network algorithm of Kohonen. The algorithm produces two-dimensional representations (maps) of speech. Input to the acoustic maps consisted of 15-component spectral vectors calculated at 9.83-msec intervals from short-time power spectra. The male and female maps were first calculated from the speech of healthy subjects and then the [a:] samples (15 successive spectral vectors) were examined on the maps. The dysphonic voices deviated from the norm both in the composition of the short-time power spectra (characterized by the dislocation of the trajectory pattern on the map) and in the stability of the spectrum during the performance (characterized by the pattern of the trajectory on the map). Rough voices were distinguished from breathy ones by their patterns on the map. With the limited speech material, an index for the degree of pathology could not be determined. A self-organized acoustic map provides an on-line visual representation of voice and speech in an easily understandable form. The method is thus suitable not only for diagnostic but also for educational and therapeutic purposes.

Adolescent↗

Speech analysis systems: an evaluation.

Performance characteristics are reviewed for seven systems marketed for acoustic speech analysis: CSpeech, CSRE, ILS-PC, Kay Elemetrics model 5500 Sona-Graph, MacSpeech Lab II, MSL, and Signalyze. The characteristics reviewed include system components, basic capabilities (signal acquisition, waveform operations, analysis, and other functions), documentation, user interface, data formats and journaling, speed and precision of spectral analysis, and speed and precision of fundamental frequency analysis. Basic capabilities are also tabulated for three recently introduced systems: the Sensimetrics SpeechStation, the Kay Elemetrics Computerized Speech Lab (CSL), and the LSI Speech Workstation. In addition to the capability and performance summaries, this article offers suggestions for continued development of speech analysis systems, particularly in data exchange, journaling, display features, spectral analysis, and fundamental frequency analysis.

Diagnosis, Computer-Assisted↗

Temporal resolution in normal-hearing and hearing-impaired listeners using frequency-modulated stimuli.

This study compares the temporal resolution of frequency-modulated sinusoids by normal-hearing and hearing-impaired subjects in a discrimination task. One signal increased linearly by 200 Hz in 50 msec. The other was identical except that its trajectory followed a series of discrete steps. Center frequencies were 500, 1000, 2000, and 4000 Hz. As the number of steps was increased, the duration of the individual steps decreased, and the subjects' discrimination performance monotonically decreased to chance. It was hypothesized that the listeners could not temporally resolve the trajectory of the step signals at short step durations. At equal sensation levels, and at equal sound pressure levels, temporal resolution was significantly reduced for the impaired subjects. The difference between groups was smaller in the equal sound pressure level condition. Performance was much poorer at 4000 Hz than at the other test frequencies in all conditions because of poorer frequency discrimination at that frequency.

Adult↗