Search PubMedSearch

Biomedical subjects

N I Durlach

Publications and source records attributed to N I Durlach.

At least 19 recordsLinked to original sources

Speaking clearly for the hard of hearing IV: Further studies of the role of speaking rate.

The contribution of reduced speaking rate to the intelligibility of "clear" speech (Picheny, Durlach, & Braida, 1985) was evaluated by adjusting the durations of speech segments (a) via nonuniform signal time-scaling, (b) by deleting and inserting pauses, and (c) by eliciting materials from a professional speaker at a wide range of speaking rates. Key words in clearly spoken nonsense sentences were substantially more intelligible than those spoken conversationally (15 points) when presented in quiet for listeners with sensorineural impairments and when presented in a noise background to listeners with normal hearing. Repeated presentation of conversational materials also improved scores (6 points). However, degradations introduced by segment-by-segment time-scaling rendered this time-scaling technique problematic as a means of converting speaking styles. Scores for key words excised from these materials and presented in isolation generally exhibited the same trends as in sentence contexts. Manipulation of pause structure reduced scores both when additional pauses were introduced into conversational sentences and when pauses were deleted from clear sentences. Key-word scores for materials produced by a professional talker were inversely correlated with speaking rate, but conversational rate scores did not approach those of clear speech for other talkers. In all experiments, listeners with normal hearing exposed to flat-spectrum background noise performed similarly to listeners with hearing loss.

Adult

A study of the tactual reception of sign language.

One of the natural methods of tactual communication in common use among individuals who are both deaf and blind is the tactual reception of sign language. In this method, the receiver (who is deaf-blind) places a hand (or hands) on the dominant (or both) hand(s) of the signer in order to receive, through the tactual sense, the various formational properties associated with signs. In the study reported here, 10 experienced deaf-blind users of either American Sign Language (ASL) or Pidgin Sign English (PSE) participated in experiments to determine their ability to receive signed materials including isolated signs and sentences. A set of 122 isolated signs was received with an average accuracy of 87% correct. The most frequent type of error made in identifying isolated signs was related to misperception of individual phonological components of signs. For presentation of signed sentences (translations of the English CID sentences into ASL or PSE), the performance of individual subjects ranged from 60-85% correct reception of key signs. Performance on sentences was relatively independent of rate of presentation in signs/sec, which covered a range of roughly 1 to 3 signs/sec. Sentence errors were accounted for primarily by deletions and phonological and semantic/syntactic substitutions. Experimental results are discussed in terms of differences in performance for isolated signs and sentences, differences in error patterns for the ASL and PSE groups, and communication rates relative to visual reception of sign language and other natural methods of tactual communication.

Adolescent

Cross-frequency interactions in the precedence effect.

This paper concerns the extent to which the precedence effect is observed when leading and lagging sounds occupy different spectral regions. Subjects, listening under headphones, were asked to match the intracranial lateral position of an acoustic pointer to that of a test stimulus composed of two binaural noise bursts with asynchronous onsets, parametrically varied frequency content, and different interaural delays. The precedence effect was measured by the degree to which the interaural delay of the matching pointer was independent of the interaural delay of the lagging noise burst in the test stimulus. The results, like those of Blauert and Divenyi [Acustica 66, 267-274 (1988)], show an asymmetric frequency effect in which the lateralization influence of a lagging high-frequency burst is almost completely suppressed by a leading low-frequency burst, whereas a lagging low-frequency burst is weighted equally with a leading high-frequency burst. This asymmetry is shown to be the result of an inherent low-frequency dominance that is seen even with simultaneous bursts. When this dominance is removed (by attenuating the low-frequency burst) the precedence effect operates with roughly equal strength both upward and downward in frequency. Within the scope of the current study (with lateralization achieved through the use of interaural time differences alone, stimuli from only two frequency bands, and only three subjects performing in all experiments), these results suggest that the precedence effect arises from a fairly central processing stage in which information is combined across frequency.

Auditory Perception

Manual discrimination of compliance using active pinch grasp: the roles of force and work cues.

In these experiments, two plates were grasped between the thumb and the index finger and squeezed together along a linear track. The force resisting the squeeze, produced by an electromechanical system under computer control, was programmed to be either constant (in the case of the force discrimination experiments) or linearly increasing (in the case of the compliance discrimination experiments) over the squeezing displacement. After completing a set of basic psychophysical experiments on compliance resolution (Experiment 1), we performed further experiments to investigate whether work and/or terminal-force cues played a role in compliance discrimination. In Experiment 2, compliance and force discrimination experiments were conducted with a roving-displacement paradigm to dissociate work cues (and terminal-force cues for the compliance experiments) from compliance and force cues, respectively. The effect of trial-by-trial feedback on response strategy was also investigated. In Experiment 3, compliance discrimination experiments were conducted with work cues totally eliminated and terminal-force cues greatly reduced. Our results suggest that people tend to use mechanical work and force cues for compliance discrimination. When work and terminal-force cues were dissociated from compliance cues, compliance resolution was poor (22%) relative to force and length resolution. When work cues were totally eliminated, performance could be predicted from terminal-force cues. A parsimonious description of all data from the compliance experiments is that subjects discriminated compliance on the basis of terminal force.

Adult

Automatic speech recognition to aid the hearing impaired: prospects for the automatic generation of cued speech.

Although great strides have been made in the development of automatic speech recognition (ASR) systems, the communication performance achievable with the output of current real-time speech recognition systems would be extremely poor relative to normal speech reception. An alternate application of ASR technology to aid the hearing impaired would derive cues from the acoustical speech signal that could be used to supplement speechreading. We report a study of highly trained receivers of Manual Cued Speech that indicates that nearly perfect reception of everyday connected speech materials can be achieved at near normal speaking rates. To understand the accuracy that might be achieved with automatically generated cues, we measured how well trained spectrogram readers and an automatic speech recognizer could assign cues for various cue systems. We then applied a recently developed model of audiovisual integration to these recognizer measurements and data on human recognition of consonant and vowel segments via speechreading to evaluate the benefit to speechreading provided by such cues. Our analysis suggests that with cues derived from current recognizers, consonant and vowel segments can be received with accuracies in excess of 80%. This level of performance is roughly equivalent to the segment reception accuracy required to account for observed levels of Manual Cued Speech reception. Current recognizers provide maximal benefit by generating only a relatively small number (three to five) of cue groups, and may not provide substantially greater aid to speechreading than simpler aids that do not incorporate discrete phonetic recognition. To provide guidance for the development of improved automatic cueing systems, we describe techniques for determining optimum cue groups for a given recognizer and speechreader, and estimate the cueing performance that might be achieved if the performance of current recognizers were improved.

Adolescent

Adjustment and discrimination measurements of the precedence effect.

A simple model to summarize the precedence effect is proposed that uses a single metric to quantify the relative dominance of the initial interaural delay over the trailing interaural delay in lateralization. This model is described and then used to relate new measurements of the precedence effect made with adjustment and discrimination paradigms. In the adjustment task, subjects matched the lateral position of an acoustic pointer to the position of a composite test stimulus made up of initial and trailing binaural noise bursts. In the discrimination procedure, subjects discriminated interaural time differences in a target noise burst in the presence of another burst either trailing or preceding the target. Experimental parameters were the delay between initial and trailing stimuli and the overall level of the stimulus. The model parameters (the metric c and the variability of lateral position judgments) were estimated from the results of the matching experiment and used to predict results of the discrimination task with good success. Finally, the observed values of the metric were compared to values derived from previous studies.

Acoustic Stimulation

Intensity perception. XIV. Intensity discrimination in listeners with sensorineural hearing loss.

Intensity discrimination of pulsed tones (also called level discrimination) was measured as a function of level in 13 listeners with sensorineural hearing impairment of primarily cochlear origin, one listener with a vestibular schwannoma, and six listeners with normal hearing. Measurements were also made in normal ears presented with masking noise spectrally shaped to produce audiograms similar to those of the cochlearly impaired listeners. For unilateral impairments, tests were made at the same frequency in the normal and impaired ears. For bilateral-sloping impairments, tests were made at different frequencies in the same ear. The normal listeners showed results similar to other data in the literature. The listener with a vestibular schwannoma showed greatly reduced intensity resolution, except at a few levels. For listeners with recruiting sensorineural impairments, the results are discussed according to the configuration of the impairment and are compared across configurations at equal SPL, equal SL, and equal loudness level. Listeners with increasing hearing losses at frequencies above the test frequency generally showed impaired resolution, especially at high levels, and less deviation from Weber's law than normal listeners. Listeners with decreasing hearing loss at frequencies above the test frequency showed nearly normal intensity-resolution functions. Whereas these trends are generally present, there are also large differences among individuals. Results obtained from normal listeners who were tested in the presence of masking noise indicate that elevated thresholds and reduced dynamic range account for some, but not all, of the effects of recruiting sensorineural impairment on intensity resolution.

Acoustic Stimulation

Analytic study of the Tadoma method: improving performance through the use of supplementary tactual displays.

Although results obtained with the Tadoma method of speechreading have set a new standard for tactual speech communication, they are nevertheless inferior to those obtained in the normal auditory domain. Speech reception through Tadoma is comparable to that of normal-hearing subjects listening to speech under adverse conditions corresponding to a speech-to-noise ratio of roughly 0 dB. The goal of the current study was to demonstrate improvements to speech reception through Tadoma through the use of supplementary tactual information, thus leading to a new standard of performance in the tactual domain. Three supplementary tactual displays were investigated: (a) an articulatory-based display of tongue contact with the hard palate; (b) a multichannel display of the short-term speech spectrum; and (c) tactual reception of Cued Speech. The ability of laboratory-trained subjects to discriminate pairs of speech segments that are highly confused through Tadoma was studied for each of these augmental displays. Generally, discrimination tests were conducted for Tadoma alone, the supplementary display alone, and Tadoma combined with the supplementary tactual display. The results indicated that the tongue-palate contact display was an effective supplement to Tadoma for improving discrimination of consonants, but that neither the tongue-palate contact display nor the short-term spectral display was highly effective in improving vowel discriminability. For both vowel and consonant stimulus pairs, discriminability was nearly perfect for the tactual reception of the manual cues associated with Cued Speech. Further experiments on the identification of speech segments were conducted for Tadoma combined with Cued Speech. The observed data for both discrimination and identification experiments are compared with the predictions of models of integration of information from separate sources.

Blindness

Development and testing of artificial low-frequency speech codes.

In a new approach to the frequency-lowering of speech, artificial codes were developed for 24 consonants (C) and 15 vowels (V) for two values of lowpass cutoff frequency F (300 and 500 Hz). Each individual phoneme was coded by a unique, nonvarying acoustic signal confined to frequencies less than or equal to F. Stimuli were created through variations in spectral content, amplitude, and duration of tonal complexes or bandpass noise. For example, plosive and fricative sounds were constructed by specifying the duration and relative amplitude of bandpass noise with various center frequencies and bandwidths, while vowels were generated through variations in the spectral shape and duration of a ten-tone harmonic complex. The ability of normal-hearing listeners to identify coded Cs and Vs in fixed-context syllables was compared to their performance on single-token sets of natural speech utterances lowpass filtered to equivalent values of F. For a set of 24 consonants in C-/a/ context, asymptotic performance on coded sounds averaged 90 percent correct for F = 500 Hz and 65 percent for F = 300 Hz, compared to 75 percent and 40 percent for lowpass filtered speech. For a set of 15 vowels in /b/-V-/t/ context, asymptotic performance on coded sounds averaged 85 percent correct for F = 500 Hz and 65 percent for F = 300 Hz, compared to 85 percent and 50 percent for lowpass filtered speech. Identification of coded signals for F = 500 Hz was also examined in CV syllables where C was selected at random from the set of 24 Cs and V was selected at random from the set of 15 Vs. Asymptotic performance of roughly 67 percent correct and 71 percent correct was obtained for C and V identification, respectively. These scores are somewhat lower than those obtained in the fixed-context experiments. Finally, results were obtained concerning the effect of token variability on the identification of lowpass filtered speech. These results indicate a systematic decrease in percent-correct score as the number of tokens representing each phoneme in the identification tests increased from one to nine.

Evaluation Studies as Topic

Manual discrimination of force using active finger motion.

In these experiments, two plates were grasped between the thumb and forefinger and squeezed together along a linear track. An electromechanical system presented a constant resistance force during the squeeze up to a predetermined location on the track, whereupon the force effectively went to infinity (simulating a wall) or to zero (simulating a cliff). The task of the subject was to discriminate between two alternative levels of the constant resistance force (a reference level and a reference-plus-increment level). Results of these experiments indicate a just noticeable difference of roughly 7% of the reference force using a one-interval paradigm with trial-by-trial feedback over the ranges 2.5 less than or equal to F0 less than or equal to 10.0 newtons, 5 less than or equal to D less than or equal to 30 mm, 45 less than or equal to S less than or equal to 125 mm, and 25 less than or equal to V less than or equal to 160 mm/sec, where F0 is the reference force, D is the distance squeezed, S is the initial fingerspan, and V is the mean velocity of the squeeze. These results, based on tests with 5 subjects, are consistent with a wide range of previous results, some of which are associated with other body surfaces and muscle systems and many of which were obtained with different psychophysical methods.

Biomechanical Phenomena

A study of the tactual and visual reception of fingerspelling.

A method of communication in frequent use among members of the deaf-blind community is the tactual reception of fingerspelling. In this method, the hand of the deaf-blind individual is placed on the hand of the sender to monitor the handshapes and movements associated with the letters of the manual alphabet. The purpose of the current study was to examine the ability of experienced deaf-blind subjects to receive fingerspelled materials, including sentences and connected text, through the tactual sense. A parallel study of the reception of fingerspelling through the visual sense was also conducted using sighted deaf subjects. For both visual and tactual reception of fingerspelled sentences, accuracy of reception was examined as a function of rate of presentation. In the tactual study, where rates were limited to those that could be produced naturally by an experienced interpreter, highly accurate reception of conversational sentence materials was observed throughout the range of naturally produced rates (i.e., 2 to 6 letters/s). In the visual study, rates in excess of those that can be produced naturally were achieved through variable-speed playback of videotapes of fingerspelled sentences. The results of this study indicate that performance varies systematically as a function of rate of presentation, with scores of 50% correct on conversational sentences obtained at rates of 12 to 16 letters/s (i.e., rates roughly double to triple normal speed). These results suggest that normal communication rates for the visual reception of fingerspelling are restricted by limitations on the rate of manual production. Although maximal rates of natural manual production of fingerspelling correspond to the presentation of a new handshape on the order of once every 150-20 ms, the data from the sped-up visual study suggest that experienced receivers of visual fingerspelling are able to receive sentences at substantially higher rates of fingerspelling (which are, in fact, comparable to communication rates for spoken English).

Adult

Analytic study of the Tadoma method: effects of hand position on segmental speech perception.

In the Tadoma method of communication, deaf-blind individuals receive speech by placing a hand on the face and neck of the talker and monitoring actions associated with speech production. Previous research has documented the speech perception, speech production, and linguistic abilities of highly experienced users of the Tadoma method. The current study was performed to gain further insight into the cues involved in the perception of speech segments through Tadoma. Small-set segmental identification experiments were conducted in which the subjects' access to various types of articulatory information was systematically varied by imposing limitations on the contact of the hand with the face. Results obtained on 3 deaf-blind, highly experienced users of Tadoma were examined in terms of percent-correct scores, information transfer, and reception of speech features for each of sixteen experimental conditions. The results were generally consistent with expectations based on the speech cues assumed to be available in the various hand positions.

Adult

Range effects in the identification of lateral position.

Experiments on the identification of interaural time and interaural amplitude differences were conducted to evaluate the effects of stimulus range on identification performance. Three stimulus sets, large range (LR), small-range center (SRC), and small-range side (SRS), were used in experiments on interaural time and amplitude identification. As expected, data for both sets of measurements show worse resolution for LR than SRC or SRS, demonstrating that the ability to distinguish between two fixed interaural differences can be strongly influenced by the total range of such differences in the stimulus set.

Auditory Perception

Analysis of a synthetic Tadoma system as a multidimensional tactile display.

The Tadoma method is a means of speech reception based on tactile monitoring of the articulatory process. A "synthetic" Tadoma system, involving an artificial face with six facial actions, has been developed as a first-order approximation to the natural Tadoma system. Experiments were conducted to explore the information-transmission characteristics of the synthetic Tadoma system in terms of the four facial movements it incorporates: upper lip in-out, lower lip in-out, lower lip up-down, and jaw up-down movements. Discrimination experiments showed that the just-noticeable difference associated with each movement is about 9% of the reference displacement. One-dimensional (1-D) absolute identification experiments produced, on the average, 1.6 bits of information transfer. Four dimensional (4-D) identification experiments produced information transfers in the range of 3-4 bits. Of the four dimensions considered, performance on the lower lip up-down movement was most affected, and performance on the jaw up-down movement was least affected, by simultaneous roving movements on the other dimensions. As a result of the interaction among the movement channels, the sum of the 1-D information transfers exceeds the 4-D information transfer. However, the sum of the 1-D information transfers obtained from tests with roving parameters is approximately equal to the 4-D information transfer (possibly exemplifying a "generalized information-transfer additivity law"). In general, both the discrimination and identification results appear unexceptional and, hence, the reception of facial movement information by itself does not appear to account for the extraordinary success of the Tadoma method.

Humans

Manual discrimination and identification of length by the finger-span method.

Experiments were conducted on length resolution for objects held between the thumb and fore-finger. The just noticeable difference in length measured in discrimination experiments is roughly 1 mm for reference lengths of 10 to 20 mm. It increases monotonically with reference length but violates Weber's law. Also, it decreases when the subject is permitted to maintain a constant finger span between trials; however, it tends to increase when the nondominant hand is used. As would be expected from studies of other stimulus dimensions in other sense modalities, resolution is considerably poorer in identification experiments than in discrimination experiments. For stimulus sets that cover a broad range (90 mm), the total information transfer is roughly 2 bits; for those that cover a relatively small range (18 mm), it is roughly 1 bit. The data are analyzed and interpreted using analysis techniques and models that have been used previously in studies of audition (e.g., Durlach & Braida, 1969).

Adult

Speaking clearly for the hard of hearing. III: An attempt to determine the contribution of speaking rate to differences in intelligibility between clear and conversational speech.

Previous studies (Picheny, Durlach, & Braida, 1985, 1986) have demonstrated that substantial intelligibility differences exist for hearing-impaired listeners for speech spoken clearly compared to speech spoken conversationally. This paper presents the results of a probe experiment intended to determine the contribution of speaking rate to the intelligibility differences. Clear sentences were processed to have the durational properties of conversational speech, and conversational sentences were processed to have the durational properties of clear speech. Intelligibility testing with hearing-impaired listeners revealed both sets of materials to be degraded after processing. However, the degradation could not be attributable to processing artifacts because reprocessing the materials to restore their original durations produced intelligibility scores close to those observed for the unprocessed materials. We conclude that the simple processing to alter the relative durations of the speech materials was not adequate to assess the contribution of speaking rate to the intelligibility differences; further studies are proposed to address this question.

Hearing Loss, Sensorineural

Tactile communication of speech: comparison of two computer-based displays.

Two methods of encoding speech for tactile displays were compared in discrimination experiments using speech segments. One display represented the short-term speech spectrum in time-swept mode and used vibration amplitude to encode spectral amplitude. The other represented the linear predictive coding (LPC)-derived vocal tract shape as a filled bar graph in which the number of active vibrators was used to encode cross sectional area. The displays were applied to the thigh via a matrix of vibrators. The vibrators were driven at 250 Hz during voiced segments, and by random noise during unvoiced segments. Overall results show a slight superiority for the spectral display in vowel discrimination. Detailed results were analyzed in terms of an articulatory description of the speech stimuli, a multidimensional scaling (MDS) analysis of confusions, and an ideal receiver analysis. The results of these analyses suggest that the detailed characteristics of the tactile patterns were only crudely discriminated.

Computer Systems