Search PubMed⌕ Search

Biomedical subjects

K N Stevens

Publications and source records attributed to K N Stevens.

At least 19 recordsLinked to original sources

An acoustical study of the fricative /s/ in the speech of individuals with dysarthria.

This paper reports on measurements of several acoustic attributes of the fricative consonant /s/ produced in word-initial position by normally speaking adults and by speakers with neuromotor dysfunctions. Several acoustic properties are evaluated: the spectrum shape of the fricative and its amplitude in relation to the following vowel, the presence or absence of voicing, the time variation of the spectrum during the fricative and in the transition to the following vowel, and the presence of inappropriate acoustic patterns preceding the /s/. Some of these properties are based on quantitative measurements of the spectrum of the /s/, and others are based on observations of the time-varying acoustic patterns in spectrograms. For the individuals with dysarthria, deviations of each of these properties from the normal range are interpreted in terms of specific deficits in the control of the speech-production system. For the most part, these parameters are highly correlated with the speakers' overall intelligibility, with the intelligibility of words containing the fricative /s/, and with perceptual ratings of the adequacy of the fricative production. The parameters that show the best correlation with intelligibility and perceptual ratings are (a) measures of deviations from normalcy in the time variation of the acoustic pattern within the consonant and at the consonant-vowel boundary and (b) the spectrum shape of the frication noise. These acoustic parameters are related to deviations in the temporal pattern of control of the articulators in producing fricative-vowel sequences and to lack of fine control of the tongue blade in achieving an appropriate target configuration for the fricative.

Adult↗

Tongue surface displacement during bilabial stops.

The goals of this study were to characterize tongue surface displacement during production of bilabial stops and to refine current estimates of vocal-tract wall impedance using direct measurements of displacement in the vocal tract during closure. In addition, evidence was obtained to test the competing claims of passive and active enlargement of the vocal tract during voicing. Tongue displacement was measured and tongue compliance was estimated in four subjects during production of /aba/ and /apa/. All subjects showed more tongue displacement during /aba/ than during /apa/, even though peak intraoral pressure is lower for /aba/. In consequence, compliance estimates were much higher for /aba/, ranging from 5.1 to 8.5 x 10(-5) cm3/dyn. Compliance values for /apa/ ranged from 0.8 to 2.3 x 10(-5) cm3/dyn for the tongue body, and 0.52 x 10(-5) for the single tongue tip point that was measured. From combined analyses of tongue displacement and intraoral pressure waveforms for one subject, it was concluded that smaller tongue displacements for /p/ than for /b/ may be due to active stiffening of the tongue during /p/, or to intentional relaxation of tongue muscles during /b/ (in conjunction with active tongue displacement during /b/).

Humans↗

Critique: articulatory-acoustic relations and their role in speech perception.

These remarks are in response to "Role of articulation in speech perception: Clues from production"ony Björn Lindblom. It is suggested that the form in which the lexicon is stored includes both segments and distinctive features, and this representation is neutral with respect to articulatory and the acoustic domains. The process by which features are determined from the sound requires that patterns of acoustic properties be identified. In developing models of speech perception, knowledge of articulatory-acoustic relations can be a guide in defining these properties, but it is not necessary for the models to assign primary status to articulation.

Humans↗

Linguistic experience alters phonetic perception in infants by 6 months of age.

Linguistic experience affects phonetic perception. However, the critical period during which experience affects perception and the mechanism responsible for these effects are unknown. This study of 6-month-old infants from two countries, the United States and Sweden, shows that exposure to a specific language in the first half year of life alters infants' phonetic perception.

Analysis of Variance↗

Acoustic and perceptual characteristics of voicing in fricatives and fricative clusters.

Several types of measurements were made to determine the acoustic characteristics that distinguish between voiced and voiceless fricatives in various phonetic environments. The selection of measurements was based on a theoretical analysis that indicated the acoustic and aerodynamic attributes at the boundaries between fricatives and vowels. As expected, glottal vibration extended over a longer time in the obstruent interval for voiced fricatives than for voiceless fricatives, and there were more extensive transitions of the first formant adjacent to voiced fricatives than for the voiceless cognates. When two fricatives with different voicing were adjacent, there were substantial modifications of these acoustic attributes, particularly for the syllable-final fricative. In some cases, these modifications leads to complete assimilation of the voicing feature. Several perceptual studies with synthetic vowel-consonant-vowel stimuli and with edited natural stimuli examined the role of consonant duration, extent and location of glottal vibration, and extent of formant transitions on the identification of the voicing characteristics of fricatives. The perceptual results were in general consistent with the acoustic observations and with expectations based on the theoretical model. The results suggest that listeners base their voicing judgments of intervocalic fricatives on an assessment of the time interval in the fricative during which there is no glottal vibration. This time interval must exceed about 60 ms if the fricative is to be judged as voiceless, except that a small correction to this threshold is applied depending on the extent to which the first-formant transitions are truncated at the consonant boundaries.

Female↗

Spectral characteristics of sound transmission in the human respiratory system.

The amplitude of sound transmission from the mouth to a site overlying the extrathoracic trachea and two sites on the right posterior chest wall over the 100-600 Hz frequency range was measured in eight healthy adult subjects. An acoustic driver and a rigid tube were employed to introduce sound into the mouths of the subjects at resting lung volume, and the transmission measurements were performed using lightweight accelerometers. Similar spectral characteristics of acceleration were observed in all of the subjects showing peaks in the transmission. These characteristics included 1) two regions of increased transmission over the frequency range of the measurements, 2) a decrease in the magnitude of acceleration of the chest wall as compared to the tracheal site of roughly 20 dB at lower frequencies, 3) a strong trend of decreasing acceleration of the chest wall with increasing frequency. These spectra agreed favorably with the predictions of a theoretical model of the acoustical properties of the respiratory system. The model suggests the primary structural determinants of a number of the observed characteristics including the importance of the lung parenchyma in sound attenuation.

Adult↗

A model of acoustic transmission in the respiratory system.

A theoretical model of sound transmission from within the respiratory tract to the chest wall due to the motion of the walls of the large airways was developed. The vocal tract, trachea, and the first five bronchial generations are represented over the frequency range from 100 to 600 Hz by an equivalent acoustic circuit. This circuit allows the estimation of the magnitude of airway wall motion in response to an acoustic perturbation at the mouth. The radiation of sound through the surrounding lung parenchyma is represented as a cylindrical wave in a homogeneous mixture of air bubbles in water. The effect of thermal losses associated with the polytropic compressions and expansions of these bubbles by the acoustic wave is included and the chest wall is represented as a massive boundary to the wave propagation. The model estimates the magnitude of acceleration over the extrathoracic trachea and at three locations on the posterior chest wall in the same vertical plane. The predicted spectral characteristics of transmission are consistent with previous experimental observations. This theoretical approach suggests that the locations of the spectral peaks are a strong function of the geometry and the wall properties of the airways, while the attenuation at higher frequencies is primarily associated with the absorption of sound in the parenchyma.

Animals↗

Calibration of ear canals for audiometry at high frequencies.

A procedure is described for determining the absolute sound pressure at the inner end of the ear canal when a sound source is coupled to the ear, for frequencies in the range 8-20 kHz. The transducer that generates the sound is coupled to the ear canal through a lossy tube, yielding a source impedance that is approximately matched to the characteristic impedance of the ear canal. A small microphone is located in the coupling tube close to the entrance to the ear canal. Calibration is carried out by measuring the response at this microphone when an impulse is applied at the transducer. To estimate the sound pressure at the medial end of the ear canal, the Fourier transform of this impulse response is corrected by an all-pole function in which the poles are estimated from the minima in this Fourier transform. Data on individual ear canals are presented in terms of gain functions relating the sound pressure at the medial end of the ear canal to the sound pressure when the coupling tube is blocked. The average gain function for a group of adult ears increases from 2 to 12 dB over the frequency range 8-20 kHz, in rough agreement with data from ear-canal models. Possible sources of error in the calibration procedure are discussed.

Acoustics↗

High-frequency audiometric assessment of a young adult population.

The hearing thresholds of 37 young adults (18-26 years) were measured at 13 frequencies (8, 9,10,...,20 kHz) using a newly developed high-frequency audiometer. All subjects were screened at 15 dB HL at the low audiometric frequencies, had tympanometry within normal limits, and had no history of significant hearing problems. The audiometer delivers sound from a driver unit to the ear canal through a lossy tube and earpiece providing a source impedance essentially equal to the characteristic impedance of the tube. A small microphone located within the earpiece is used to measure the response of the ear canal when an impulse is applied at the driver unit. From this response, a gain function is calculated relating the equivalent sound-pressure level of the source to the SPL at the medial end of the ear canal. For the subjects tested, this gain function showed a gradual increase from 2 to 12 dB over the frequency range. The standard deviation of the gain function was about 2.5 dB across subjects in the lower frequency region (8-14 kHz) and about 4 dB at the higher frequencies. Cross modes and poor fit of the earpiece to the ear canal prevented accurate calibration for some subjects at the highest frequencies. The average SPL at threshold was 23 dB at 8 kHz, 30 dB at 12 kHz, and 87 dB at 18 kHz. Despite the homogeneous nature of the sample, the younger subjects in the sample had reliably better thresholds than the older subjects. Repeated measurements of threshold over an interval as long as 1 month showed a standard deviation of 2.5 dB at the lower frequencies (8-14 kHz) and 4.5 dB at the higher frequencies.

Adolescent↗

Acoustic and perceptual correlates of the non-nasal--nasal distinction for vowels.

For each of five vowels [i e a o u] following [t], a continuum from non-nasal to nasal was synthesized. Nasalization was introduced by inserting a pole-zero pair in the vicinity of the first formant in an all-pole transfer function. The frequencies and spacing of the pole and zero were systematically varied to change the degree of nasalization. The selection of stimulus parameters was determined from acoustic theory and the results of pilot experiments. The stimuli were presented for identification and discrimination to listeners whose language included a non-nasal--nasal vowel opposition (Gujarati, Hindi, and Bengali) and to American listeners. There were no significant differences between language groups in the 50% crossover points of the identification functions. Some vowels were more influenced by range and context effects than were others. The language groups showed some differences in the shape of the discrimination functions for some vowels. On the basis of the results, it is postulated that (1) there is a basic acoustic property of nasality, independent of the vowel, to which the auditory system responds in a distinctive way regardless of language background; and (2) there are one or more additional acoustic properties that may be used to various degrees in different languages to enhance the contrast between a nasal vowel and its non-nasal congener. A proposed candidate for the basic acoustic property is a measure of the degree of prominence of the spectral peak in the vicinity of the first formant. Additional secondary properties include shifts in the center of gravity of the low-frequency spectral prominence, leading to a change in perceived vowel height, and changes in overall spectral balance.

Humans↗

On some issues in the pursuit of acoustic invariance in speech: a reply to Lisker.

This paper presents an alternative view of acoustic invariance in speech to that discussed by Lisker [J. Acoust. Soc. Am. 77, 1199-1202 (1985)]. Three points are considered--the minimal unit for acoustic invariance, the level of linguistic representation over which this unit operates, and the role that acoustic invariance plays in speech and language. Our position emphasizes the role of phonetic features in a theory of acoustic invariance, and we propose a series of working hypotheses to guide research in this area.

Humans↗

Effect of burst amplitude on the perception of stop consonant place of articulation.

We have examined the effects of the relative amplitude of the release burst on perception of the place of articulation of utterance-initial voiceless and voiced stop consonants. The amplitude of the burst, which occurs within the first 10-15 ms following consonant release, was systematically varied in 5-dB steps from -10 to +10 dB relative to a "normal" burst amplitude for two labial-to-alveolar synthetic speech continua--one comprising voiceless stops and the other, voiced stops. The distribution of spectral energy in the bursts for the labial and alveolar stops at the ends of the continuum was consistent with the spectrum shapes observed in natural utterances, and intermediate shapes were used for intermediate stimuli on the continuum. The results of identification tests with these stimuli showed that the relative amplitude of the burst significantly affected the perception of the place of articulation of both voiceless and voiced stops, but the effect was greater for the former than the latter. The results are consistent with a view that two basic properties contribute to the labial-alveolar distinction in English. One of these is determined by the time course of the change in amplitude in the high-frequency range (above 2500 Hz) in the few tens of ms following consonantal release, and the other is determined by the frequencies of spectral peaks associated with the second and third formants in relation to the first formant.

Humans↗

Perceptual invariance and onset spectra for stop consonants in different vowel environments.

A series of listening tests with brief synthetic consonant-vowel syllables was carried out to determine whether the initial part of a syllable can provide cues to place of articulation for voiced stop consonants independent of the remainder of the syllable. The data show that stimuli as short as 10-20 ms sampled from the onset of a consonant-vowel syllable, can be reliably identified for consonantal place of articulation, whether the second and higher formants contain moving or straight transitions and whether or not an initial burst is present. In most instances, these brief stimuli also contain sufficient information for vowel indentification. Stimulus continua in which formant transitions ranged from values appropriate to [b], [d], [g] in various vowel environments, and in which stimulus durations were 20 and 46 ms, yielded categorical labeling functions with a few exceptions. These results are consistent with a theory of speech perception in which consonant place of articulation is cued by invariant properties derived from the spectrum sampled in a 10-20 ms time window adjacent to consonantal onset or offset.

Adult↗

Acoustic correlates of some phonetic categories.

Some of the acoustic properties that distinguish one speech sound from another are reviewed. The point of view is that the auditory system responds to sound with different acoustic properties in distinctive ways, and that these special responses play an important role in selection and classification of the inventory of sounds that are used in language. Examples of several of these acoustic properties are discussed and illustrated, including the presence or absence of rapid spectrum change, abruptness and amplitude change, voicing and aspiration, and gross spectral properties relating to place of articulation for consonant and vowels.

Humans↗

Acoustic invariance in speech production: evidence from measurements of the spectral characteristics of stop consonants.

On the basis of theoretical considerations and the results of experiments with synthetic consonant-vowel syllables, it has been hypothesized that the short-time spectrum sampled at the onset of a stop consonant should exhibit gross properties that uniquely specify the consonantal place of articulation independent of the following vowel. The aim of this paper is to test this hypothesis by measuring the spectrum sampled at the onsets and offsets of a large number of consonant-vowel (CV) and vowel-consonant (VC) syllables containing both voiced and voiceless stops produced by several speakers. Templates were devised in an attempt to capture three classes of spectral shapes: diffuse-rising, diffuse-falling, and compact, corresponding to alveolar, labial, and velar consonants, respectively. Spectra were derived from the utterances by sampling at the consonantal release of CV syllables and at the implosion and burst release of VC syllables, and these spectra (smoothed by a linear prediction algorithm) were matched against the templates. It was found that about 85% of the spectra at initial consonant release and at final burst release were correctly classified by the templates, although there was some variability across vowel contexts. The spectra sampled at the implosion were not consistently classified. A preliminary examination of spectra sampled at the release of nasal consonants in CV syllables showed a somewhat lower accuracy of classification by the same templates. Overall, the results support an hypothesis that, in natural speech, the acoustic characteristics of stop consonants, specified in terms of the gross spectral shape sampled at the discontinuity in the acoustic signal, show invariant properties independent of the adjacent vowel or of the voicing characteristics of the consonant. The implication is that the auditory system is endowed with detectors that are sensitive to these kinds of gross spectral shapes, and that the existence of these detectors helps the infant to organize the sounds of speech into their natural classes.

Humans↗

Invariant cues for place of articulation in stop consonants.

In a series of experiments, identification responses for place of articulation were obtained for synthetic stop consonants in consonant-vowel syllables with different vowels. The acoustic attributes of the consonants were systematically manipulated, the selection of stimulus characteristics being guided in part by theoretical considerations concerning the expected properties of the sound generated in the vocal tract as place of articulation is varied. Several stimulus series were generated with and without noise bursts at the onset, and with and without formant transitions following consonantal release. Stimuli with transitions only, and with bursts plus transitions, were consistently classified according to place of articulation, whereas stimuli with bursts only and no transitions were not consistently identified. The acoustic attributes of the stimuli were examined to determine whether invariant properties characterized each place of atriculation independent of vowel context. It was determined that the gross shape of the spectrum sampled at the consonantal release showed a distinctive shape for each place of articulation: a prominent midfrequency spectral peak for velars, a diffuse-rising spectrum for alveolars, and a diffuse-falling spectrum for labials. These attributes are evident for stimuli containing transitions only, but are enhanced by the presence of noise bursts at the onset.

Acoustic Stimulation↗