Search PubMed⌕ Search

Biomedical subjects

D H Whalen

Publications and source records attributed to D H Whalen.

At least 19 recordsLinked to original sources

Phonetic processing areas revealed by sinewave speech and acoustically similar non-speech.

The neural substrates underlying speech perception are still not well understood. Previously, we found dissociation of speech and nonspeech processing at the earliest cortical level (AI), using speech and nonspeech complexity dimensions. Acoustic differences between speech and nonspeech stimuli in imaging studies, however, confound the search for linguistic-phonetic regions. Presently, we used sinewave speech (SWsp) and nonspeech (SWnon), which replace speech formants with sinewave tones, in order to match acoustic spectral and temporal complexity while contrasting phonetics. Chord progressions (CP) were used to remove the effects of auditory coherence and object processing. Twelve normal RH volunteers were scanned with fMRI while listening to SWsp, SWnon, CP, and a baseline condition arranged in blocks. Only two brain regions, in bilateral superior temporal sulcus, extending more posteriorly on the left, were found to prefer the SWsp condition after accounting for acoustic modulation and coherence effects. Two regions responded preferentially to the more frequency-modulated stimuli, including one that overlapped the right temporal phonetic area and another in the left angular gyrus far from the phonetic area. These findings are proposed to form the basis for the two subtypes of auditory word deafness. Several brain regions, including auditory and non-auditory areas, preferred the coherent auditory stimuli and are likely involved in auditory object recognition. The design of the current study allowed for separation of acoustic spectrotemporal, object recognition, and phonetic effects resulting in distinct and overlapping components.

Adolescent↗

Differentiation of speech and nonspeech processing within primary auditory cortex.

Primary auditory cortex (PAC), located in Heschl's gyrus (HG), is the earliest cortical level at which sounds are processed. Standard theories of speech perception assume that signal components are given a representation in PAC which are then matched to speech templates in auditory association cortex. An alternative holds that speech activates a specialized system in cortex that does not use the primitives of PAC. Functional magnetic resonance imaging revealed different brain activation patterns in listening to speech and nonspeech sounds across different levels of complexity. Sensitivity to speech was observed in association cortex, as expected. Further, activation in HG increased with increasing levels of complexity with added fundamentals for both nonspeech and speech stimuli, but only for nonspeech when separate sources (release bursts/fricative noises or their nonspeech analogs) were added. These results are consistent with the existence of a specialized speech system which bypasses more typical processes at the earliest cortical level.

Acoustic Stimulation↗

A sex difference in visual influence on heard speech.

Reports of sex differences in language processing are inconsistent and are thought to vary by task type and difficulty. In two experiments, we investigated a sex difference in visual influence onheard speech (the McGurk effect). First, incongruent consonant-vowel stimuli were presented where the visual portion of the signal was brief (100 msec) or full (temporally equivalent to the auditory). Second, to determine whether men and women differed in their ability to extract visual speech information from these brief stimuli, the same stimuli were presented to new participants with an additional visual-only (lipread) condition. In both experiments, women showed a significantly greater visual influence on heard speech than did men for the brief visual stimuli. No sex differences for the full stimuli or in the ability to lipread were found. These findings indicate that the more challenging brief visual stimuli elicit sex differences in the processing of audiovisual speech.

Adult↗

The Haskins optically corrected ultrasound system (HOCUS).

The tongue is critical in the production of speech, yet its nature has made it difficult to measure. Not only does its ability to attain complex shapes make it difficult to track, it is also largely hidden from view during speech. The present article describes a new combination of optical tracking and ultrasound imaging that allows for a noninvasive, real-time view of most of the tongue surface during running speech. The optical system (Optotrak) tracks the location of external structures in 3-dimensional space using infrared emitting diodes (IREDs). By tracking 3 or more IREDs on the head and a similar number on an ultrasound transceiver, the transduced image of the tongue can be corrected for the motion of both the head and the transceiver and thus be represented relative to the hard structures of the vocal tract. If structural magnetic resonance images of the speaker are available, they may allow the estimation of the location of the rear pharyngeal wall as well. This new technique is contrasted with other currently available options for imaging the tongue. It promises to provide high-quality, relatively low-cost imaging of most of the tongue surface during fairly unconstrained speech.

Humans↗

Perception of pitch location within a speaker's F0 range.

Fundamental frequency (F0) is used for many purposes in speech, but its linguistic significance is based on its relation to the speaker's range, not its absolute value. While it may be that listeners can gauge a specific pitch relative to a speaker's range by recognizing it from experience, whether they can do the same for an unfamiliar voice is an open question. The present experiment explored that question. Twenty native speakers of English (10 male, 10 female) produced the vowel /a/ with a spoken (not sung) voice quality at varying pitches within their own ranges. Listeners then judged, without familiarization or context, where each isolated F0 lay within each speaker's range. Correlations were high both for the entire range (0.721) and for the range minus the extremes (0.609). Correlations were somewhat higher when the F0s were related to the range of all the speakers, either separated by sex (0.830) or pooled (0.848), but several factors discussed here may help account for this pattern. Regardless, the present data provide strong support for the hypothesis that listeners are able to locate an F0 reliably within a range without external context or prior exposure to a speaker's voice.

Adult↗

Vowel production and perception: hyperarticulation without a hyperspace effect.

The ability of speakers to exaggerate speech sounds ("hyperarticulation") has led to the theory that the targets themselves must be hyperspace hyperarticulated. Johnson, Flemming, and Wright (1993) found that perceptual "best exemplar" choices for vowels were more speech extreme than listeners' own productions. Our first experiment, using their procedure, only partially replicated their results. Low vowels vowel perception showed a higher F1, consistent with hyperspace. Front vowels also showed more frontness in F2, but back vowels were less extreme ("hypoarticulated") on F2. Our second experiment used an identification and rating of each stimulus, yielding similar results of a smaller magnitude. Our results indicate that the perceptual space is calibrated to a particular (synthetic) vowel space, which is not related straightforwardly to the speakers' spaces. The original hyperspace hypothesis can be attributed to the methodology which led to extreme judgments and of the fronting of back vowels in California English. The present results indicate that no such hypothesis is needed. Vowel targets are measurable from an individual's productions, and the individual's perception of other speakers (even synthetic ones) is based on information about the vocal tract and dialect of the speaker.

Adult↗

Posterior pharyngeal wall position in the production of speech.

The posterior pharyngeal wall has been assumed to be stationary during speech. The present study examines this assumption in order to assess whether midsagittal widths in the pharyngeal region can be inferred from measurements of the anterior pharyngeal wall. Midsagittal magnetic resonance images and X-ray images were examined to determine whether the posterior pharyngeal wall from the upper oropharynx to the upper laryngopharynx shows anterior movement that can be attributed to variables in speech: vowel quality in both English and Japanese; vowels versus consonants as classes of speech sounds; sustained versus dynamically produced speech; and isolated words versus sentences. Measurements were made of the distance between the anterior portion of the vertebral body and the pharyngeal wall. The first measurement was on a line traversing the junction between the dens and the body of the second cervical vertebra (C2). The next three measurements were on lines at the inferior borders of the bodies of C2, C3, and C4. The measurements showed very little movement of the posterior pharyngeal wall, none of it attributable to speech variables. Therefore, the position of the posterior pharyngeal wall in this region can be eliminated as a variable, and the anterior portion of the pharynx alone can be used to estimate vocal cavities.

Adult↗

Parametrically dissociating speech and nonspeech perception in the brain using fMRI.

Candidate brain regions constituting a neural network for preattentive phonetic perception were identified with fMRI and multivariate multiple regression of imaging data. Stimuli contrasted along speech/nonspeech, acoustic, or phonetic complexity (three levels each) and natural/synthetic dimensions. Seven distributed brain regions' activity correlated with speech and speech complexity dimensions, including five left-sided foci [posterior superior temporal gyrus (STG), angular gyrus, ventral occipitotemporal cortex, inferior/posterior supramarginal gyrus, and middle frontal gyrus (MFG)] and two right-sided foci (posterior STG and anterior insula). Only the left MFG discriminated natural and synthetic speech. The data also supported a parallel rather than serial model of auditory speech and nonspeech perception.

Adult↗

Predicting midsagittal pharynx shape from tongue position during vowel production.

The shape of the pharynx has a large effect on the acoustics of vowels, but direct measurement of this part of the vocal tract is difficult. The present study examines the efficacy of inferring midsagittal pharynx shape from the position of the tongue, which is much more amenable to measurement. Midsagittal magnetic resonance (MR) images were obtained for multiple repetitions of 11 static English vowels spoken by two subjects (one male and one female). From these, midsagittal widths were measured at approximately 3-mm intervals along the entire vocal tract. A regression analysis was then used to assess whether the pharyngeal widths could be predicted from the locations and width measurements for four positions on the tongue, namely, those likely to be the locations of a receiver coil for an electromagnetometer system. Predictability was quite high throughout the vocal tract (multiple r> 0.9), except for the extreme ends (i.e., larynx and lips) and small decreases for the male subject in the uvula region. The residuals from this analysis showed that the accuracy of predictions was generally quite high, with 89.2% of errors being less than 2 mm. The extremes of the vocal tract, where the resolution of the MRI was poorer, accounted for much of the error. For languages like English, which do not use advanced tongue root (ATR) distinctively, the midsagittal pharynx shape of static vowels can be predicted with high accuracy.

Female↗

Exploring the relationship of inspiration duration to utterance duration.

Previous work has indicated that there may be a positive relationship between the duration and extent of inspiration and the length of an upcoming utterance. However, none of that work has uniquely implied a role of planning. We attempted to avoid some of the alternative explanations by forcing subjects to utter single sentences ranging in length from 5 to 82 syllables (mean of 27), after inspiring fully and then expiring down to a set level before uttering the sentence. For all 3 subjects, there was a significant positive relationship between utterance length and inspiration duration, regardless of whether inspiration was measured physiologically or acoustically. The 2 subjects with the higher correlations in the articulatory measures also expended air more quickly during the shorter sentences than longer ones, while the other subject had no correlation with exhalation rate. Complexity of the sentence, calculated as the number of clauses in the sentence, did not affect inspiration duration. The individual differences need further investigation, but there is a positive correlation between the duration of the sentence to be said and the inspiration before it when the speaker is required to read sentences while using only one breath.

Humans↗

Limits on phonetic integration in duplex perception.

The telling fact about duplex perception is that listeners integrate into a unitary phonetic percept signals that are coherent from a phonetic point of view, even though the signals are, on purely auditory grounds, separate sources. Here we explore the limits on the integration of a sinusoidal consonant cue (the F3 transition for [da] vs. [ga]) with the resonances of the remainder of the syllable. Perceiving duplexly, listeners hear the whistle of the sinusoid, but also the [da] and [ga] for which the sinusoid provides the critical information. In the first experiment, phonetic integration was significantly reduced, but not to zero, by a precursor that extended the transition cue forward in time so that it started 50 msec before the cue. The effect was the same above and below the duplexity threshold (the intensity of sinusoid in the combined pattern at which the whistle was just barely audible). In the second experiment, integration was reduced once again by the precursor, and also, but only below the duplexity threshold, by harmonics of the cues that were simultaneous with it. The third experiment showed that the simultaneous harmonics reduced phonetic integration only by serving as distractors while also permitting the conclusion that the precursor produced its effects by making the cue part of a coherent and competing auditory pattern, and so "capturing" it. The fourth experiment supported this interpretation by showing that for some subjects the amount of capture was reduced when the capturing tone was itself captured by being made part of a tonal complex. The results support the assumption that the independent phonetic system will integrate across disparate sources according to the cohesive power of that system as measured against the evidence for separate sources.

Adult↗

The effects of breath sounds on the perception of synthetic speech.

When preparing to speak, talkers typically take a breath. The perceptual effect of adding naturally produced breath intake sounds to synthetic speech was examined. In experiment 1, subjects were better at transcribing synthesized sentences that were preceded by a breath sound than those that were not, in addition to the improvement due to practice that is typically found with synthetic speech. Experiment 2 found that replacing the breath with the spectrally similar sound of rustling leaves had no effect on the accuracy. Experiment 3 had breaths before randomly selected sentences. Only the practice effect was significant, though there was a tendency for sentences with the breath sounds to be remembered better. In experiment 4, we tested whether the appropriateness of the breath sound to the sentence size (relatively short or long) affected the use of the breath sound. Appropriateness had no effect, perhaps because the range of sentence durations was too small. Experiment 5 replicated experiment 1 but used leaf sounds rather than silence in the nonbreath sentences. The presence of breath was again found to aid recall. Overall, the current results indicate that adding the breath intake sound to synthetic sentences improves listeners' ability to recall those sentences.

Female↗

Intrinsic F0 of vowels in the babbling of 6-, 9-, and 12-month-old French- and English-learning infants.

In every language so far examined, high vowels such as [i] and [u] tend to have higher fundamental frequencies (F0s) than low vowels such as [a]. This intrinsic F0 effect (IF0) has been found in the speech of children at various stages of development, except in the one previous study of babbling. The present study is based on a larger set of utterances from more subjects (six French- and six English-learning infants), at the ages 6, 9, and 12 months. It is found, instead, that IF0 appears even in babbling. There is no indication in these data of a developmental trend for the effect, and no indication of a difference due to the target language. These results support the claim that IF0 is an automatic consequence of producing vowels.

Child Development↗

FO gives voicing information even with unambiguous voice onset times.

The voiced/voiceless distinction for English utterance-initial stop consonants is primarily realized as differences in the voice onset time (VOT), which is largely signaled by the time between the stop burst and the onset of voicing. The voicing of stops has also been shown to affect the vowel's FO after release, with voiceless stops being associated with higher FO. When the VOT is ambiguous, these FO "perturbations" have been shown to affect voicing judgments. This is to be expected of what can be considered a redundant feature, that is, that it should carry a distinction in cases where the primary feature is neutralized. However, when the voicing judgments were made as quickly as possible, an inappropriate FO was found to slow response time even for unambiguous VOTs. This was true both of FO contours and level FO differences. These results reinforce the plausibility of tonogenesis, and they add further weight to the claim that listeners make full use of the signal given to them, even when overt labeling would seem to indicate otherwise.

Audiometry↗

Information for Mandarin tones in the amplitude contour and in brief segments.

While the tones of Mandarin are conveyed mainly by the F0 contour, they also differ consistently in duration and in amplitude contour. The contribution of these factors was examined by using signal-correlated noise stimuli, in which natural speech is manipulated so that it has no F0 or formant structure but retains its original amplitude contour and duration. Tones 2, 3 and 4 were perceptible from just the amplitude contour, even when duration was not also a cue. In two further experiments, the location of the critical information for the tones during the course of the syllable was examined by extracting small segments from each part of the original syllable. Tones 2 and 3 were often confused with each other, and segments which did not have much F0 change were most often heard as Tone 1. There were, though, also cases in which a low, unchanging pitch was heard as Tone 3, indicating a partial effect of register even in Mandarin. F0 was positively correlated with amplitude, even when both were computed on a pitch period basis. Taken together, the results show that Mandarin tones are realized in more than just the F0 pattern, that amplitude contours can be used by listeners as cues for tone identification, and that not every portion of the F0 pattern unambiguously indicates the original tone.

Adult↗

Intonational differences between the reduplicative babbling of French- and English-learning infants.

The two- and three-syllable reduplicative babbling of five French-learning and five English-learning infants (0;5 to 1;1) was examined in two ways for intonational differences. The first measure was a categorization into one of five categories (RISING, FALLING, RISE-FALL, FALL-RISE, LEVEL) by expert listeners. The second was the fundamental frequency (F0) from the early, middle and late portion of each syllable. Both measures showed significant differences between the two language groups. 65% of the utterances from both groups were classified as either rising of falling. For the French children, these were divided equally into the rising and the falling categories, while 75% of those utterances for the English children were judged to have falling intonation. Proportions of the other three categories were not significantly different by language environment. In both languages, though, three-syllable utterances were more likely to have a complex contour than two-syllable ones. Analysis of the F0 patterns confirmed the perceptual assessment. Several aspects of the target languages help explain these intonational differences in prelinguistic babbling.

Data Interpretation, Statistical↗

Perception of the English /s/-/integral of/ distinction relies on fricative noises and transitions, not on brief spectral slices.

A series of experiments compared two approaches to fricative identification, spectral template matching and articulatory dynamics. Natural-speech /s/ and /integral of/ noises from fricative-vowel or vowel-fricative syllables were cross spliced so that "hybrid" noises started out as either /s/ or /integral of/ and ended up with the other fricative noise in varying proportions. With both initial and final fricatives, listener judgments most often agreed with the longer part of the noise even when spectral templates would predict the other category. Also, the vocalic formant transitions contributed to the judgment. In another experiment, open transcriptions by four expert listeners similarly showed that all the cues were used; there were also some instances of nonspeech percepts that would be predicted by gestural models. One further experiment had subjects identify two fricatives from hybrid noises between two vocalic segments. When the order of the noises differed from the order of the transitions, the perceived ordering of the fricatives was often the reverse of the order of the noise segments. Taken together with previous results, these experiments indicate that listeners take the whole fricative noise, as well as the transitions, into account in fricative identification.

Adult↗