Search PubMed⌕ Search

Biomedical subjects

R N Ohde

Publications and source records attributed to R N Ohde.

At least 19 recordsLinked to original sources

Temporal processing in the aging auditory system.

Measures of monaural temporal processing and binaural sensitivity were obtained from 12 young (mean age = 26.1 years) and 12 elderly (mean age = 70.9 years) adults with clinically normal hearing (pure-tone thresholds < or = 20 dB HL from 250 to 6000 Hz). Monaural temporal processing was measured by gap detection thresholds. Binaural sensitivity was measured by interaural time difference (ITD) thresholds. Gap and ITD thresholds were obtained at three sound levels (4, 8, or 16 dB above individual threshold). Subjects were also tested on two measures of speech perception, a masking level difference (MLD) task, and a syllable identification/discrimination task that included phonemes varying in voice onset time (VOT). Elderly listeners displayed poorer monaural temporal analysis (higher gap detection thresholds) and poorer binaural processing (higher ITD thresholds) at all sound levels. There were significant interactions between age and sound level, indicating that the age difference was larger at lower stimulus levels. Gap detection performance was found to correlate significantly with performance on the ITD task for young, but not elderly adult listeners. Elderly listeners also performed more poorly than younger listeners on both speech measures; however, there was no significant correlation between psychoacoustic and speech measures of temporal processing. Findings suggest that age-related factors other than peripheral hearing loss contribute to temporal processing deficits of elderly listeners.

Adult↗

Stop-consonant and vowel perception in 3- and 4-year-old children.

Recent research on 5- to 11-year-old children's perception of stop consonants and vowels indicates that they can generally identify these sounds with relatively high accuracy from short duration stimulus onsets [Ohde et al., J. Acoust. Soc. Am. 97, 3800-3812 (1995); Ohde et al., J. Acoust. Soc. Am. 100, 3813-3824 (1996)]. The purpose of the current experiments was to determine if younger children, aged 3-4 years, can also recover consonant and vowel features from stimulus onsets. Ten adults, ten 3-year olds, and ten 4-year-olds listened to synthesized syllables composed of combinations of [b d g] and [i u a]. The synthesis parameters included manipulations of the following stimulus variables: formant transition (moving or straight), noise burst (present or absent), and voicing duration (10, 30, or 46 ms). Developmental effects were found for the perception of both stop consonants and vowels. In general, adults identified these sounds at a significantly higher level than children, and perception by 4-year-olds was significantly better than 3-year-olds. A developmental effect of dynamic formant motion was obtained, but it was limited to only the [g] stop consonant. Stimulus duration affected the children's perception of vowels indicating that they may utilize additional auditory information to a much greater extent than adults. The results support the importance of information in stimulus onsets for syllable identification, and developmental changes in sensitivity to these cues for consonant and vowel perception.

Adult↗

A developmental study of vowel perception from brief synthetic consonant-vowel syllables.

The purpose of this study was to assess the perceptual role of brief synthetic consonant-vowel syllables as cues for vowel perception in children and adults. Nine types of consonant-vowel syllables comprised of the stops [b d g] followed by the vowels [i a u] were synthesized. Stimuli were generated with durations of 10, 30, or 46 ms, and with or without formant transition motion. Eight children at each of five age levels (5, 6, 7, 9, and 11 years) and a control group of eight adults were trained to identify each vowel in a three-alternative forced-choice (3AFC) paradigm. The results showed that children and adults extracted vowel information at a generally high level from stimuli as brief as 10 ms. For many stimuli, there was little or no difference between the performance of children and adults. However, developmental effects were observed. First, the accuracy of vowel perception was more influenced by the consonant context for children than for adults. Whereas perception was similar across age levels for stimuli in the alveolar context, the youngest children perceived vowels in the labial and velar contexts at significantly lower levels than adults. Second, children were more affected by variations in stimulus duration than were adults. This finding was particularly prominent for the syllable [ga], where the dependency on duration decreased with age in a nearly linear fashion. These findings are discussed in relation to current hypotheses of vowel perception in adults, and hypotheses of speech perception development.

Adult↗

The effect of segment duration on the perceptual integration of nasals for adult and child speech.

It has been hypothesized that the acoustic properties within a temporal domain of 10 to 30 ms of boundaries between speech sounds contain significant information on the phonetic features of segments, and that these cues are perceptually integrated by the auditory system [Stevens, Phonetic Linguistics: Essays in Honor of Peter Ladefoged (Academic, London, 1985)]. The purpose of the current research was to examine the effects of stimulus duration adjacent to speech sound boundaries on the perceptual integration of place of articulation of nasals before and after disruption of the abrupt changes in spectra between the murmur and transition. In experiment I, three children, aged 3, 5, and 7 years, and an adult female and male produced consonant-vowel (CV) syllables consisting of [m] and [n] in four vowel contexts, [i ae u a]. Approximately 25-ms segments of the murmur and vowel transition adjacent to the speech sound boundary were digitally removed from these productions. Intervals of silence ranging from 0 to 2000 ms, which can potentially perturb integration processes, were inserted between these segments. The stimuli were then presented to adult listeners for the identification of the nasal. The main findings revealed a consistent decline in identification with gap durations up to 150 ms across speakers and vowel context. However, the adult labial feature was resistant to perceptual change as a function of gap duration. This result appeared to relate to formant transition duration, and not to response bias. In experiment II, stimuli with durations shorter than those in experiment I were further analyzed for adult speakers. The main finding was a quantification of the acoustic segment duration needed for perceptual integration of the murmur and vowel transition. Across both experiments, the results reveal a decline in the identification of both alveolar and labial nasals within a time interval mediated by short-term auditory memory, and that the duration of the acoustic segment needed for perceptual integration is longer for [n] than [m].

Adult↗

Stimulus uncertainty and speaker normalization processes in the perception of nasal consonants.

Previous research has noted a reduction in perceptual identification performance when the speaker varies from stimulus to stimulus and has interpreted this finding as an effect of a normalization process that compensates for variability in the physical content of the speech signal. The purpose of the present investigation was to examine whether the physical variability introduced by this experimental design can result in general stimulus uncertainty effects that extend beyond the realm of traditional delineations of normalization processes. The stimuli were short segments taken from nasal consonant + vowel syllables produced by 1 male adult, 1 female adult, and 2 children. Segments of 25 and 50 ms duration were edited from the nasal murmur and the onset of the vowel. The stimuli were ordered according to four variability conditions and presented to listeners for place of articulation identification. The results showed that identification was significantly reduced when variability affected either speaker identity or segment type. Further analyses revealed that segment variability impaired perception of all segment types approximately equally, but that the 50-ms vowel segments were selectively spared in the speaker variability condition. These findings indicate that general uncertainty effects should be considered in speech perception experiments, and that dynamic properties of speech are particularly important in the perceptual compensation for speaker variability.

Adult↗

A developmental study of the perception of onset spectra for stop consonants in different vowel environments.

The importance of different acoustic properties for the perception of place of articulation in prevocalic stop consonants was investigated from a developmental perspective. Eight adults and eight children in each of the age groups, 5, 6, 7, 9, and 11 years, listened to synthesized syllables comprised of all combinations of [b d g] and [i a]. The synthesis parameters were adapted from Blumstein and Stevens [J. Acoust. Soc. Am. 67, 648-662 (1980)], and included manipulations of the following stimulus variables: formant transitions (moving or straight), noise burst (present or absent), and voicing duration (10 or 46 ms). Identification performance was high for all age groups across most stimulus types. Formant transition motion generally was not necessary for accurate identification, and there was no difference between age groups in terms of the perceptual weight placed on this cue. Furthermore, the results did not support the salience of duration as a developmental cue to place of articulation. The presence of a burst improved identification for the velar and alveolar places of articulation for all age groups, but was particularly important for the 11-year-olds and adults. These findings indicate that children, by age 5, do not rely on dynamic formant motion any more than adults do, and that the ability to integrate acoustic cues across regions of spectral change shows developmental patterns.

Adult↗

The role of short-term and long-term auditory storage in processing spectral relations for adult and child speech.

The processes involved in the perception of spectral change between the nasal murmur and the vocalic transition for speakers of different ages were assessed before and after disruption of the variation in spectra between these elements. Three children, aged 3, 5, and 7, and an adult female and male produced consonant-vowel (CV) syllables consisting of either [m] or [n] followed by [i] or [u]. In one condition (spectrally noncontiguous), the acoustic information surrounding the region of spectral change was digitally removed and in another condition (spectrally contiguous) this portion of the signal was retained. In both of these conditions, intervals of silence ranging from 0 to 2000 ms were inserted between 50-ms segments of murmur and vocalic transition. These gap duration conditions were then presented to adult listeners for the identification of the nasal. Across speakers, the results for the spectral contiguous condition support a primary mechanism in the perception of spectral relations that is mediated by processes within short-term auditory memory, but the results for the spectral noncontiguous condition revealed little consistent support for either short-term or long-term memory processes.

Age Factors↗

The development of the perception of cues to the [m]-[n] distinction in CV syllables.

The contribution of the nasal murmur and vocalic formant transition to the perception of the [m]-[n] distinction by adult listeners was investigated for speakers of different ages. Children, aged 3, 5, and 7, and an adult female and male produced consonant-vowel (CV) syllables consisting of either [m] or [n] and followed by [i ae u a]. Three productions of each syllable were computer edited into the following segments: (1) full murmur; (2) 50-ms murmur preceding release; (3) 25-ms murmur preceding release; (4) 25-ms murmur preceding release +25-ms transition following release; (5) 25-ms transition following release; (6) 50-ms transition following release; and (7) full transition+vowel. The results indicate that both the murmur and transition provide cues to place of articulation, but that the latter property is more prominent in perception than the former across speaker age. The salience of the murmur and vocalic transition cues was greater for adults than children indicating a developmental progression of the encoding of gestures associated with these properties. Although the simultaneous presence of murmur and vocalic transition cues surrounding the point of spectral discontinuity improved perception of place of articulation across speakers, there was evidence of a developmental progression of this property also. For speakers of all ages, as segment duration decreased, consistent decrements in identification of place of articulation occurred only for transition stimuli. The murmur+transition was the most salient cue supporting the importance of spectral discontinuities and/or relational properties in production and perception particularly in the acquisition of sound features.

Age Factors↗

Effect of relative amplitude of frication on perception of place of articulation.

The amplitude of frication relative to vowel onset amplitude in the F3 and F5 formant frequency regions was manipulated for the synthetic fricative contrasts /s/-/integral of/ and /s/-/theta/, respectively. The influence of this relative amplitude manipulation on listeners' perception of place of articulation was tested by (1) varying the duration of frication from 30 to 140 ms, (2) pairing the frication noise with different vowels /i a u/, (3) placing formant transitions in conflict with relative amplitude, and (4) holding relative amplitude constant within a continuum while varying formant transitions and the amplitudes of spectral regions where relative amplitude was not manipulated. To determine if listeners were using absolute spectral cues or relative amplitude comparisons between frication and vowel for fricative identification, the frication and vowel were separated by (1) presenting the frication in isolation, and (2) inserting a gap of silence between the frication and vowel. The results showed that relative amplitude was perceived across vowel context and frication duration, and overrode context-dependent formant transition cues. The findings for temporal separations between the frication and vowel suggest that short-term memory processes may dominate the mediation of the relative-amplitude comparison. However, the overall results indicate that relative amplitude is only a component of spectral prominence, which is comprised of a primary frication spectral peak and a secondary frication/vowel peak comparison.

Female↗

Spectral and duration properties of front vowels as cues to final stop-consonant voicing.

The perception of voicing in final velar stop consonants was investigated by systematically varying vowel duration, change in offset frequency of the final first formant (F1) transition, and rate of frequency change in the final F1 transition for several vowel contexts. Consonant-vowel-consonant (CVC) continua were synthesized for each of three vowels, [i,I,ae], which represent a range of relatively low to relatively high-F1 steady-state values. Subjects responded to the stimuli under both an open- and closed-response condition. Results of the study show that both vowel duration and F1 offset properties influence perception of final consonant voicing, with the salience of the F1 offset property higher for vowels with high-F1 steady-state frequencies than low-F1 steady-state frequencies, and the opposite occurring for the vowel duration property. When F1 onset and offset frequencies were controlled, rate of the F1 transition change had inconsistent and minimal effects on perception of final consonant voicing. Thus the findings suggest that it is the termination value of the F1 offset transition rather than rate and/or duration of frequency change, which cues voicing in final velar stop consonants during the transition period preceding closure.

Adult↗

Frequency discrimination ability and stop-consonant identification in normally hearing and hearing-impaired subjects.

Identification of place of articulation in the synthesized syllables /bi/, /di/, and /gi/ was examined in three groups of listeners: (a) normal hearers, (b) subjects with high-frequency sensorineural hearing loss, and (c) normally hearing subjects listening in noise. Stimuli with an appropriate second formant (F2) transition (moving-F2 stimuli) were compared with stimuli in which F2 was constant (straight-F2 stimuli) to examine the importance of the F2 transition in stop-consonant perception. For straight-F2 stimuli, burst spectrum and F2 frequency were appropriate for the syllable involved. Syllable duration also was a variable, with formant durations of 10, 19, 28, and 44 ms employed. All subjects' identification performance improved as stimulus duration increased. The groups were equivalent in terms of their identification of /di/ and /gi/ syllables, whereas the hearing-impaired and noise-masked normal listeners showed impaired performance for /bi/, particularly for the straight-F2 version. No difference in performance among groups was seen for /di/ and /gi/ stimuli for moving-F2 and straight-F2 versions. Second-formant frequency discrimination measures suggested that subjects' discrimination abilities were not acute enough to take advantage of the formant transition in the /di/ and /gi/ stimuli.

Adult↗

Relationship between the discrimination of (w-r) and (t-d) continua and the identification of distorted (r).

The purpose of this study was to assess the extent to which listeners can perceive intraphonemic differences. In Experiment 1, subjects identified synthesized acoustic tokens of child-like speech that varied in second and third formant (F2 and F3) onset frequencies as (w), (r), or distorted (r) in two conditions: (a) with and without feedback of the group response choices, and (b) before and after training to identify the best examples of (w), (r), and distorted (r) based on their identification in the first condition. The results were: (a) some subjects consistently identified distorted (r) above criterion, and (b) feedback was more effective in increasing distorted (r) identification than was training. In Experiment 2, the same subjects participated in discrimination tasks using stimuli from a synthesized child (w-r) continuum that varied in F2 and F3 onsets and from a synthesized adult (t-d) continuum that varied in preconsonantal vowel duration. The results were: (a) perception was not categorical for both continua, (b) little relation was found between distorted-(r) identification and measures of (w-r) discrimination, and (c) a high and significant correlation was found between identification of distorted (r) and within-(d) discrimination. In Experiment 3, different subjects identified the child manifold stimuli and discriminated stimuli in a synthesized child (w-r) continuum and in a synthesized adult (t-d) continuum. The results were: (a) neither (w-r) or (t-d) perception was categorical although the former came closer than the latter in terms of individual subject performance, (b) there was a high and significant correlation between distorted-(r) identification and within-(r) discrimination of (w-r) stimuli, and (c) there were high and significant correlations between distorted-(r) identification and mean, cross-category boundary, and within-(t) discrimination of (t-d) stimuli.

Adult↗

Perceptual categorization and consistency of synthesized (r-w) continua by adults, normal children and (r)-misarticulating children.

The purpose of this study was to determine if children who misarticulate (r) differ from normal children and adults in the perception of sound features that are produced correctly and incorrectly. Children with normal articulation, children who produced (r) misarticulations, and adults listened to synthesized child and adult (r-w) continua in two separate sessions, and to an adult (b-w) control continuum in one session. Perception was evaluated on the basis of measures of phonetic boundary location and the consistency of response to each stimulus in a continuum. The (r)-misarticulating children were found to be significantly less consistent than child and adult controls in responding to the (r-w) stimuli. Moreover, consistency scores were significantly higher for the adult continuum than for the child continuum. The performance of children was different from that of adults. Due to inconsistent performance, boundaries could not be computed for (r)-misarticulating children, but it was found that the boundaries for children in the control group were closer to the (r)-end of the continuum than those for adults. In the case of the (b-w) continuum, it was found that (r)-misarticulating children were significantly less consistent than adults. The phonetic boundaries of children were significantly closer to the (b)-end of the continuum than the boundary for adults. Thus, the results reveal that variability in stimulus response was influenced primarily by the productive ability of the subjects, whereas differences in stimulus categorization were influenced by the age of the subjects. The perceptual variability was most clearly reflected by responses to stimuli produced incorrectly, whereas categorization differences extended to sounds produced correctly.

Adult↗

Revisiting stop-consonant perception for two-formant stimuli.

The purpose of this study was to reexamine the factors leading to stop-consonant perception for consonant-vowel (CV) stimuli with just two formants over a range of vowels, under both an open- and closed-response condition. Five two-formant CV stimulus continua were synthesized, each covering a range of second-formant (F2) starting frequencies, for vowels corresponding roughly to [i,I,ae,u,a]. In addition, for the [I] and [a] continua, the duration of the first-formant (F1) transition was systematically varied. Three main findings emerged. First, criterion-level labial and alveolar responses were obtained for those stimuli with substantial F2 transitions. Second, for some stimuli, increases in the duration of the F1 transition increased velar responses to criterion level. Third, the response paradigm had a substantial influence on stop-consonant perception across all vowel continua. The results support a model of stop-consonant perception that includes spectral and time-varying spectral properties as integral components of analysis.

Adult↗

Effect of formant transition rate on the differentiation of synthesized child and adult (w) and (r) sounds.

The purpose of this research was to assess the perceptual effects of a range of second (F2) and third (F3) formant transition rates that occur naturally in the production of /w/ and /r/ by children and adults. Synthesized CV continua that varied in the second (F2) and third (F3) formant onset frequencies between values appropriate for /w/ and /r/ were used as stimuli. Subjects participated in four experimental conditions that involved changing the rate of transition of either F2 or F3 by varying the duration of the transition between the glide onset and an /eI/ vowel nucleus for child and adult stimuli. In each condition, the transition rates of the /w/-endpoint stimulus, the /r/-endpoint stimulus, and at least one of the midpoint stimuli were varied across values appropriate for /w/ and /r/. Mean ratings of the stimuli were compared to test the predictions that /w/ perception increases with increased F2 transition rate and increases with decreased F3 transition rate. The results were as follows: (a) one of the 10 comparisons for the child F2 stimuli was significant, but it involved a change opposite to the predicted direction; (b) four of the six comparisons for the child F3 stimuli were significant, but they involved changes opposite to the predicted direction; (c) 14 comparisons of adult F2 stimuli were significant, but 12 of these 14 comparisons involved changes opposite to the predicted direction; and (d) four of the six comparisons for the adult F3 stimuli were significant, but none of them involved changes in the predicted direction. Only one of all the comparisons involved a significant change in rating between the /w/ and /r/ categories.

Child↗

Fundamental frequency correlates of stop consonant voicing and vowel quality in the speech of preadolescent children.

Fundamental frequency (F0) and voice onset time (VOT) were measured in utterances containing voiceless aspirated [ph, th, kh], voiceless unaspirated [sp, st, sk], and voiced [b, d, g] stop consonants produced in the context of [i, e, u, o, a] by 8- to 9-year-old subjects. The results revealed that VOT reliably differentiated voiceless aspirated from voiceless unaspirated and voiced stops, whereas F0 significantly contrasted voiced with voiceless aspirated and unaspirated stops, except for the first glottal period, where voiceless unaspirated stops contrasted with the other two categories. Fundamental frequency consistently differentiated vowel height in alveolar and velar stop consonant environments only. In comparing the results of these children and of adults, it was observed that the acoustic correlates of stop consonant voicing and vowel quality were different not only in absolute values, but also in terms of variability. Further analyses suggested that children were more variable in production due to inconsistency in achieving specific targets. The findings also suggest that, of the acoustic correlates of the voicing feature, the primary distinction of VOT is strongly developed by 8-9 years of age, whereas the secondary distinction of F0 is still in an emerging state.

Adult↗

Effect of formant frequency onset variation on the differentiation of synthesized /w/ and /r/ sounds.

The purpose of this study was to assess the use of psychophysical transformations for analyzing the differentiation of /w/ and /r/ sounds of children and adults. Stimuli from Adult and Child manifolds, consisting of 25 synthesized /Cej/-type utterances with different F2 and F3 onset frequencies, were presented in random order to eight naive subjects. Subjects rated the stimuli on a four-point scale between good /r/ and good /w/. Correlations between mel transformations and Bark transformations of the F3-F2 differences among the stimuli and their percent /r/ responses were close to or greater than .90. Predictions of percent /r/ responses derived from regression analyses based on mel transformations and Bark transformations of F3-F2 differences among stimuli indicated that some sounds identified as /w/ for /r/ substitutions could be differentiated from /w/ sounds. The category boundaries between /r/ and /w/ were estimated to be 5.0 Bark for adult stimuli and 5.7 Bark for child stimuli.

Adult↗

Fundamental frequency as an acoustic correlate of stop consonant voicing.

Fundamental frequency (F0) and voice onset time (VOT) were measured in utterances containing voiceless aspirated /ph,th,kh/, voiceless unaspirated /sp,st,sk/, and voiced /b,d,g/ stop consonants. Although VOT was very similar for voiceless unaspirated and voiced stops, F0 contours were nearly identical for voiceless unaspirated and voiceless aspirated stops, and both types of voiceless stops were associated with significantly higher F0 values than were voiced stops. The F0 contours in all context were generally falling; the data do not support a simple rise-fall dichotomy in F0 at voicing onset as an invariant acoustic correlate of the voicing feature. The variations in F0 as a function of voicing appear to be best accounted for by vocal cord tension rather than aerodynamic influences. Moreover, the results are consistent with physiological data showing that the position of the hyoid bone and the height of the larynx influence the absolute value of F0.

Humans↗