Search PubMed⌕ Search

Biomedical subjects

D B Pisoni

Publications and source records attributed to D B Pisoni.

At least 73 records · Page 4Linked to original sources

Phonological priming in auditory word recognition.

Cohort theory, developed by Marslen-Wilson and Welsh (1978), proposes that a "cohort" of all the words beginning with a particular sound sequence will be activated during the initial stage of the word recognition process. We used a priming technique to test specific predictions regarding cohort activation in three experiments. In each experiment, subjects identified target words embedded in noise at different signal-to-noise ratios. The target words were either presented in isolation or preceded by a prime item that shared phonological information with the target. In Experiment 1, primes and targets were English words that shared zero, one, two, three, or all phonemes from the beginning of the word. In Experiment 2, nonword primes preceded word targets and shared initial phonemes. In Experiment 3, word primes and word targets shared phonemes from the end of a word. Evidence of reliable phonological priming was observed in all three experiments. The results of the first two experiments support the assumption of activation of lexical candidates based on word-initial information, as proposed in cohort theory. However, the results of the third experiment, which showed increased probability of correctly identifying targets that shared phonemes from the end of words, did not support the predictions derived from the theory. The findings are discussed in terms of current models of auditory word recognition and recent approaches to spoken-language understanding.

Cues↗

Speech perception: some new directions in research and theory.

The perception of speech is one of the most fascinating attributes of human behavior; both the auditory periphery and higher centers help define the parameters of sound perception. In this paper some of the fundamental perceptual problems facing speech sciences are described. The paper focuses on several of the new directions speech perception research is taking to solve these problems. Recent developments suggest that major breakthroughs in research and theory will soon be possible. The current study of segmentation, invariance, and normalization are described. The paper summarizes some of the new techniques used to understand auditory perception of speech signals and their linguistic significance to the human listener.

Attention↗

Infant discrimination of two- and five-formant voiced stop consonants differing in place of articulation.

According to recent theoretical accounts of place of articulation perception, global, invariant properties of the stop CV syllable onset spectrum serve as primary, innate cues to place of articulation, whereas contextually variable formant transitions constitute secondary, learned cues. By this view, one might expect that young infants would find the discrimination of place of articulation contrasts signaled by formant transition differences more difficult than those cued by gross spectral differences. Using an operant head-turning paradigm, we found that 6-month-old infants were able to discriminate two-formant stimuli contrasting in place of articulation as well as they did five-formant + burst stimuli. Apparently, neither the global properties of the onset spectrum nor simply the additional acoustic information contained in the five-formant + burst stimuli afford the infant any advantage in the discrimination task. Rather, formant transition information provides a sufficient basis for discriminating place of articulation differences.

Child Language↗

Identification and discrimination of rise time: is it categorical or noncategorical?

Previous studies have reported that rise time of sawtooth waveforms may be discriminated in either a categorical-like manner under some experimental conditions or according to Weber's law under other conditions. In the present experiments, rise time discrimination was examined with two experimental procedures: the traditional labeling and ABX tasks used in speech perception studies and an adaptive tracking procedure used in psychophysical studies. Rise time varied from 0 to 80 ms in 10-ms intervals for sawtooth signals of 1-s duration. Discrimination functions for subjects who simply discriminated the signals on any basis whatsoever as well as functions for subjects who practiced labeling the endpoint stimuli as " pluck " and "bow" before ABX discrimination were not categorical in the ABX task. In the adaptive tracking procedure, the Weber fraction obtained from the jnds of rise time was found to be a constant above 20-ms rise time. The results from the two discrimination paradigms were then compared by predicting a jnd for rise time from the ABX discrimination data by reference to the underlying psychometric function. Using this method of analysis, discrimination results from previous studies were shown to be quite similar to the discrimination results observed in this study. Taken together the results demonstrate clearly that rise time discrimination of sawtooth signals follows predictions derived from Weber's law.

Adult↗

Recognition of speech spectrograms.

The performance of eight naive observers in learning to identify speech spectrograms was studied over a 2-month period. Single tokens from a 50-word phonetically balanced (PB) list were recorded by several talkers and displayed on a Spectraphonics Speech Spectrographic Display system. Identification testing occurred immediately after daily training sessions. After approximately 20 h of training, naive subjects correctly identified the 50 PB words from a single talker over 95% of the time. Generalization tests with the same words were then carried out with different tokens from the original talker, new tokens from another male talker, a female talker, and finally, a synthetic talker. The generalization results for these talkers showed recognition performance at 91%, 76%, 76%, and 48%, respectively. Finally, generalization tests with a novel set of PB words produced by the original talker were also carried out to examine in detail the perceptual strategies and visual features that subjects abstracted from the training set. Our results demonstrate that even without formal training in phonetics or acoustics naive observers can learn to identify visual displays of speech at very high levels of accuracy. Analysis of subjects' performance in a verbal protocol task demonstrated that they rely on salient visual correlates of many phonetic features in speech.

Communication Methods, Total↗

Infants' discrimination of the duration of a rapid spectrum change in nonspeech signals.

Two-month-old infants discriminated complex sinusoidal patterns that varied in the duration of their initial frequency transitions. Discrimination of these nonspeech sinusoidal patterns was a function of both the duration of the transitions and the total duration of the stimulus pattern. This contextual effect was observed even though the information specifying stimulus duration occurred after the transitional information. These findings parallel those observed with infants for perception of synthetic speech stimuli. Specialized speech processing capacities are thus not required to account for infants' sensitivity to contextual effects in acoustic signals, whether speech or nonspeech.

Humans↗

Coding of the speech spectrum in three time-varying sinusoids.

Recent perceptual experiments with normal adult listeners show that phonetic information can readily be conveyed by sinewave replicas of speech signals. These tonal patterns are made of three sinusoids set equal in frequency and amplitude to the respective peaks of the first three formants of natural-speech utterances. Unlike natural and most synthetic speech, the spectrum of sinusoidal patterns contains neither harmonics nor broadband formants, and is identified as grossly unnatural in voice timbre. Despite this drastic recoding of the short-time speech spectrum, listeners perceive the phonetic content if the temporal properties of spectrum variation are preserved. These observations suggest that phonetic perception may depend on properties of coherent spectrum variation, a second-order property of the acoustic signal, rather than any particular set of acoustic elements present in speech signals.

Humans↗

Vibrotactile identification of vowels.

The ability of subjects to identify vowels in vibrotactile transformations of consonant-vowel syllables was measured for two types of displays: a spectral display (frequency by intensity), and a vocal tract area function display (vocal tract location by cross-sectional area). Both displays were presented to the fingertip via the tactile display of the Optacon transducer. In the first experiments the spectral display was effective for identifying vowels in /b/V/ context when as many as 24 or as few as eight spectral channels were presented to the skin. However, performance fell when the 12- and 8-channel displays were reduced in size to occupy 1/2 or 1/3 of the 24-row tactile matrix. The effect of reducing the size of the display was greater when the spectrum was represented as a solid histogram ("filled" patterns) than when it was represented as a simple spectral contour ("unfilled" patterns). Spatial masking within the filled pattern was postulated as the cause for this decline in performance. Another experiment measured the utility of the spectral display when the syllables were produced by multiple speakers. The resulting increase in response confusions was primarily attributable to variations in the tactile patterns caused by differences in vocal tract resonances among the speakers. The final experiment found an area function display to be inferior to the spectral display for identification of vowels. The results demonstrate that a two-dimensional spectral display is worthy of further development as a basic vibrotactile display for speech.

Communication Devices for People with Disabilities↗

Perception of static and dynamic acoustic cues to place of articulation in initial stop consonants.

Two recent accounts of the acoustic cues which specify place of articulation in syllable-initial stop consonants claim that they are located in the initial portions of the CV waveform and are context-free. Stevens and Blumstein [J. Acoust. Soc. Am. 64, 1358-1368 (1978)] have described the perceptually relevant spectral properties of these cues as static, while Kewley-Port [J. Acoust. Soc. Am. 73, 322-335 (1983)] describes these cues as dynamic. Three perceptual experiments were conducted to test predictions derived from these accounts. Experiment 1 confirmed that acoustic cues for place of articulation are located in the initial 20-40 ms of natural stop-vowel syllables. Next, short synthetic CV's modeled after natural syllables were generated using either a digital, parallel-resonance synthesizer in experiment 2 or linear prediction synthesis in experiment 3. One set of synthetic stimuli preserved the static spectral properties proposed by Stevens and Blumstein. Another set of synthetic stimuli preserved the dynamic properties suggested by Kewley-Port. Listeners in both experiments identified place of articulation significantly better from stimuli which preserved dynamic acoustic properties than from those based on static onset spectra. Evidently, the dynamic structure of the initial stop-vowel articulatory gesture can be preserved in context-free acoustic cues which listeners use to identify place of articulation.

Communication Devices for People with Disabilities↗

Some effects of laboratory training on identification and discrimination of voicing contrasts in stop consonants.

For many years there has been a consensus that early linguistic experience exerts a profound and often permanent effect on the perceptual abilities underlying the identification and discrimination of stop consonants. It has also been concluded that selective modification of the perception of stop consonants cannot be accomplished easily and quickly in the laboratory with simple discrimination training techniques. In the present article we report the results of three experiments that examined the perception of a three-way voicing contrast by naive monolingual speakers of English. Laboratory training procedures were implemented with a small computer in a real-time environment to examine the perception of voiced, voiceless unaspirated, and voiceless aspirated stops differing in voice onset time. Three perceptual categories were present for most subjects after only a few minutes of exposure to the novel contrast. Subsequent perceptual tests revealed reliable and consistent labeling and categorical-like discrimination functions for all three voicing categories, even though one of the contrasts is not phonologically distinctive in English. The present results demonstrate that the perceptual mechanisms used by adults in categorizing stop consonants can be modified easily with simple laboratory techniques in a short period of time.

Adult↗

Speech perception without traditional speech cues.

A three-tone sinusoidal replica of a naturally produced utterance was identified by listeners, despite the readily apparent unnatural speech quality of the signal. The time-varying properties of these highly artificial acoustic signals are apparently sufficient to support perception of the linguistic message in the absence of traditional acoustic cues for phonetic segments.

Auditory Perception↗