Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sound Spectrography”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,747 records · Page 97Linked to original sources

Infinite-impulse-response models of the head-related transfer function.

Head-related transfer functions (HRTFs) measured from human subjects were approximated using infinite-impulse-response (IIR) filter models. Models were restricted to rational transfer functions (plus simple delays) so that specific models are characterized by the locations of poles and zeros in the complex plane. The all-pole case (with no nontrivial zeros) is treated first using the theory of linear prediction. Then the general pole-zero model is derived using a weighted-least-squares (WLS) formulation of the modified least-squares problem proposed by Kalman (1958). Both estimation algorithms are based on solutions of sets of linear equations and result in efficient computational schemes to find low-order model HRTFs. The validity of each of these two low-order models was assessed in psychophysical experiments. Specifically, a four-interval, two-alternative, forced-choice paradigm was used to test the discriminability of virtual stimuli constructed from empirical and model HRTFs for corresponding locations. For these experiments, the stimuli were 80 ms, noise tokens generated from a wideband noise generator. Results show that sounds synthesized through model HRTFs were indistinguishable from sounds synthesized from original HRTF measurements for the majority of positions tested. The advantages of the techniques described here are the computational efficiencies achieved for low-order IIR models. Properties of the all-pole and pole-zero estimators are discussed in the context of low-order HRTF representations, and implications for basic and applied contexts are considered.

Acoustics↗

Low-frequency whale and seismic airgun sounds recorded in the mid-Atlantic Ocean.

Beginning in February 1999, an array of six autonomous hydrophones was moored near the Mid-Atlantic Ridge (35 degrees N-15 degrees N, 50 degrees W-33 degrees W). Two years of data were reviewed for whale vocalizations by visually examining spectrograms. Four distinct sounds were detected that are believed to be of biological origin: (1) a two-part low-frequency moan at roughly 18 Hz lasting 25 s which has previously been attributed to blue whales (Balaenoptera musculus); (2) series of short pulses approximately 18 s apart centered at 22 Hz, which are likely produced by fin whales (B. physalus); (3) series of short, pulsive sounds at 30 Hz and above and approximately 1 s apart that resemble sounds attributed to minke whales (B. acutorostrata); and (4) downswept, pulsive sounds above 30 Hz that are likely from baleen whales. Vocalizations were detected most often in the winter, and blue- and fin whale sounds were detected most often on the northern hydrophones. Sounds from seismic airguns were recorded frequently, particularly during summer, from locations over 3000 km from this array. Whales were detected by these hydrophones despite its location in a very remote part of the Atlantic Ocean that has traditionally been difficult to survey.

Animal Communication↗

Accurate analysis of multitone signals using a DFT.

Optimum data windows make it possible to determine accurately the amplitude, phase, and frequency of one or more tones (sinusoidal components) in a signal. Procedures presented in this paper can be applied to noisy signals, signals having moderate nonstationarity, and tones close in frequency. They are relevant to many areas of acoustics where sounds are quasistationary. Among these are acoustic probes transmitted through media and natural sounds, such as animal vocalization, speech, and music. The paper includes criteria for multitone FFT block design and an example of application to sound transmission in the atmosphere.

Acoustics↗

The influence of duration and level on human sound localization.

The localization of sounds in the vertical plane (elevation) deteriorates for short-duration wideband sounds at moderate to high intensities. The effect is described by a systematic decrease of the elevation gain (slope of stimulus-response relation) at short sound durations. Two hypotheses have been proposed to explain this finding. Either the sound localization system integrates over a time window that is too short to accurately extract the spectral localization cues (neural integration hypothesis), or the effect results from cochlear saturation at high intensities (adaptation hypothesis). While the neural integration model predicts that elevation gain is independent of sound level, the adaptation hypothesis holds that low elevation gains for short-duration sounds are only obtained at high intensities. Here, these predictions are tested over a larger range of stimulus parameters than has been done so far. Subjects responded with rapid head movements to noise bursts in the two-dimensional frontal space. Stimulus durations ranged from 3 to 100 ms; sound levels from 26 to 73 dB SPL. Results show that the elevation gain decreases for short noise bursts at all sound levels, a finding that supports the integration model. On the other hand, the short-duration gain also decreases at high sound levels, which is in line with the adaptation hypothesis. The finding that elevation gain was a nonmonotonic function of sound level for all sound durations, however, is predicted by neither model. It is concluded that both mechanisms underlie the elevation gain effect and a conceptual model is proposed to reconcile these findings.

Adult↗

Effect of number of masking talkers and auditory priming on informational masking in speech recognition.

Three experiments investigated factors that influence the creation of and release from informational masking in speech recognition. The target stimuli were nonsense sentences spoken by a female talker. In experiment 1 the masker was a mixture of three, four, six, or ten female talkers, all reciting similar nonsense sentences. Listeners' recognition performance was measured with both target and masker presented from a front loudspeaker (F-F) or with a masker presented from two loudspeakers, with the right leading the front by 4 ms (F-RF). In the latter condition the target and masker appear to be from different locations. This aids recognition performance for one- and two-talker maskers, but not for noise. As the number of masking talkers increased to ten, the improvement in the F-RF condition diminished, but did not disappear. The second experiment investigated whether hearing a preview (prime) of the target sentence before it was presented in masking improved recognition for the last key word, which was not included in the prime. Marked improvements occurred only for the F-F condition with two-talker masking, not for continuous noise or F-RF two-talker masking. The third experiment found that the benefit of priming in the F-F condition was maintained if the prime sentence was spoken by a different talker or even if it was printed and read silently. These results suggest that informational masking can be overcome by factors that improve listeners' auditory attention toward the target.

Adult↗

Measuring sperm whales from their clicks: stability of interpulse intervals and validation that they indicate whale length.

Multiple pulses can often be distinguished in the clicks of sperm whales (Physeter macrocephalus). Norris and Harvey [in Animal Orientation and Navigation, NASA SP-262 (1972), pp. 397-417] proposed that this results from reflections within the head, and thus that interpulse interval (IPI) is an indicator of head length, and by extrapolation, total length. For this idea to hold, IPIs must be stable within individuals, but differ systematically among individuals of different size. IPI stability was examined in photographically identified individuals recorded repeatedly over different dives, days, and years. IPI variation among dives in a single day and days in a single year was statistically significant, although small in magnitude (it would change total length estimates by <3%). As expected, IPIs varied significantly among individuals. Most individuals showed significant increases in IPIs over several years, suggesting growth. Mean total lengths calculated from published IPI regressions were 13.1 to 16.1 m, longer than photogrammetric estimates of the same whales (12.3 to 15.3 m). These discrepancies probably arise from the paucity of large (12-16 m) whales in data used in published regressions. A new regression is offered for this size range.

Acoustics↗

Dominance of missing fundamental versus spectrally cued pitch: individual differences for complex tones with unresolved harmonics.

In a two-alternative, forced-choice experiment, subjects had to compare the pitches of two sounds, A and B. Each sound was composed of four successive harmonics of a fundamental frequency between 100 to 250 Hz, added in cosine or Schröder phase. The harmonic frequencies of A were lower than those of B; the missing fundamental frequency of A was higher than that of B. The dominance of the missing fundamental versus the spectrally cued pitch--a pitch percept corresponding to spectral components--was measured as a function of nA, the lowest harmonic in A. The pitch percept is dominated by the missing fundamental if the harmonics are resolved (nA<7). If the harmonics become unresolved and are added in Schröder phase, the dominance shifts to a spectrally cued pitch (7 20). For others, the transition was in the realm of partly resolved harmonics. This shows that the temporal envelope modulation of stimuli with only four unresolved harmonics can give a relatively clear fundamental pitch percept. However, this percept varies considerably among subjects.

Acoustic Stimulation↗

Echolocation signals of dusky dolphins (Lagenorhynchus obscurus) in Kaikoura, New Zealand.

An array of four hydrophones arranged in a symmetrical star configuration was used to measure the echolocation signals of the dusky dolphin (Lagenorhynchus obscurus) near the Kaikoura Peninsula, New Zealand. Most of the echolocation signals had bi-modal frequency spectra with a low-frequency peak between 40 and 50 kHz and a high-frequency peak between 80 and 110 kHz. The low-frequency peak was dominant when the source level was low and the high frequency peak dominated when the source level was high. The center frequencies in the dusky broadband echolocation signals are among the highest of dolphins measured in the field. Peak-to-peak source levels as high as 210 dB re 1 microPa were measured, although the average was much lower in value. The levels of the echolocation signals are about 9-12 dB lower than for the larger white-beaked dolphin (Lagenorhynchus albirostris) which belongs to the same genus but is over twice as heavy as the dusky dolphins. The source level varied in amplitude approximately as a function of the one-way transmission loss for signals traveling from the animals to the array. The wave form and spectrum of the echolocation signals were similar to those of other dolphins measured in the field.

Animals↗

Potential sound production by a deep-sea fish.

Swimbladder sonic muscles of deep-sea fishes were described over 35 years ago. Until now, no recordings of probable deep-sea fish sounds have been published. A sound likely produced by a deep-sea fish has been isolated and localized from an analysis of acoustic recordings made at the AUTEC test range in the Tongue of the Ocean, Bahamas, from four deep-sea hydrophones. This sound is typical of a fish sound in that it is pulsed and relatively low frequency (800-1000 Hz). Using time-of-arrival differences, the sound was localized to 548-696-m depth, where the bottom was 1620 m. The ability to localize this sound in real-time on the hydrophone range provides a great advantage for being able to identify the sound-producer using a remotely operated vehicle.

Air Sacs↗

Auditory steady-state responses reveal amplitude modulation gap detection thresholds.

Auditory evoked magnetic fields were recorded from the left hemisphere of healthy subjects using a 37-channel magnetometer while stimulating the right ear with 40-Hz amplitude modulated (AM) tone-bursts with 500-Hz carrier frequency in order to study the time-courses of amplitude and phase of auditory steady-state responses (ASSRs). The stimulus duration of 300 ms and the duration of the silent periods (3-300 ms) between succeeding stimuli were chosen to address the question whether the time-course of the ASSR can reflect both temporal integration and temporal resolution in the central auditory processing. Long lasting perturbations of the ASSR were found after gaps in the AM sound, even for gaps of short duration. These were interpreted as evidences for an auditory reset mechanism. Concomitant psycho-acoustical tests corroborated that gap durations perturbing the ASSR were in the same range as the threshold for AM gap detection. Magnetic source localizations estimated the ASSR sources in the primary auditory cortex, suggesting that the processing of temporal structures in the sound is performed at or below the cortical level.

Acoustic Stimulation↗

Independent component analysis for automatic note extraction from musical trills.

The method of principal component analysis, which is based on second-order statistics (or linear independence), has long been used for redundancy reduction of audio data. The more recent technique of independent component analysis, enforcing much stricter statistical criteria based on higher-order statistical independence, is introduced and shown to be far superior in separating independent musical sources. This theory has been applied to piano trills and a database of trill rates was assembled from experiments with a computer-driven piano, recordings of a professional pianist, and commercially available compact disks. The method of independent component analysis has thus been shown to be an outstanding, effective means of automatically extracting interesting musical information from a sea of redundant data.

Acoustics↗

Therelationship between professional operatic soprano voice and high range spectral energy.

Operatic sopranos need to be audible over an orchestra yet they are not considered to possess a singer's formant. As in other voice types, some singers are more successful than others at being heard and so this work investigated the frequency range of the singer's formant between 2000 and 4000 Hz to consider the question of extra energy in this range. Such energy would give an advantage over an orchestra, so the aims were to ascertain what levels of excess energy there might be and look at any relationship between extra energy levels and performance level. The voices of six operatic sopranos (national and international standard) were recorded performing vowel and song tasks and subsequently analyzed acoustically. Measures taken from vowel data were compared with song task data to assess the consistency of the approaches. Comparisons were also made with regard to two conditions of intended projection (maximal and comfortable), two song tasks (anthem and aria), two recording environments (studio and anechoic room), and between subjects. Ranking the singers from highest energy result to lowest showed the consistency of the results from both vowel and song methods and correlated reasonably well with the performance level of the subjects. The use of formant tuning is considered and examined.

Adult↗

A neural network model of the articulatory-acoustic forward mapping trained on recordings of articulatory parameters.

Three neural network models were trained on the forward mapping from articulatory positions to acoustic outputs for a single speaker of the Edinburgh multi-channel articulatory speech database. The model parameters (i.e., connection weights) were learned via the backpropagation of error signals generated by the difference between acoustic outputs of the models, and their acoustic targets. Efficacy of the trained models was assessed by subjecting the models' acoustic outputs to speech intelligibility tests. The results of these tests showed that enough phonetic information was captured by the models to support rates of word identification as high as 84%, approaching an identification rate of 92% for the actual target stimuli. These forward models could serve as one component of a data-driven articulatory synthesizer. The models also provide the first step toward building a model of spoken word acquisition and phonological development trained on real speech.

Adult↗

The effect of loading on disturbance sounds of the Atlantic croaker Micropogonius undulatus: air versus water.

Physiological work on fish sound production may require exposure of the swimbladder to air, which will change its loading (radiation mass and resistance) and could affect parameters of emitted sounds. This issue was examined in Atlantic croaker Micropogonius chromis by recording sounds from the same individuals in air and water. Although sonograms appear relatively similar in both cases, pulse duration is longer because of decreased damping, and sharpness of tuning (Q factor) is higher in water. However, pulse repetition rate and dominant frequency are unaffected. With appropriate caution it is suggested that sounds recorded in air can provide a useful tool in understanding the function of various swimbladder adaptations and provide reasonable approximation of natural sounds. Further, they provide an avenue for experimentally manipulating the sonic system, which can reveal details of its function not available from intact fish underwater.

Acoustics↗

Contrasting monaural and interaural spectral cues for human sound localization.

A human psychoacoustical experiment is described that investigates the role of the monaural and interaural spectral cues in human sound localization. In particular, it focuses on the relative contribution of the monaural versus the interaural spectral cues towards resolving directions within a cone of confusion (i.e., directions with similar interaural time and level difference cues) in the auditory localization process. Broadband stimuli were presented in virtual space from 76 roughly equidistant locations around the listener. In the experimental conditions, a "false" flat spectrum was presented at the left eardrum. The sound spectrum at the right eardrum was then adjusted so that either the true right monaural spectrum or the true interaural spectrum was preserved. In both cases, the overall interaural time difference and overall interaural level difference were maintained at their natural values. With these virtual sound stimuli, the sound localization performance of four human subjects was examined. The localization performance results indicate that neither the preserved interaural spectral difference cue nor the preserved right monaural spectral cue was sufficient to maintain accurate elevation judgments in the presence of a flat monaural spectrum at the left eardrum. An explanation for the localization results is given in terms of the relative spectral information available for resolving directions within a cone of confusion.

Acoustic Stimulation↗

Blind deconvolution of audio-frequency signals using the self-deconvolving data restoration algorithm.

A signal processing algorithm has been developed in which a filter function is extracted from degraded data through mathematical operations. The filter function can be used to restore much of the degraded content of the data through use of a deconvolution process. The operation can be performed without prior knowledge of the detection system, a technique known as blind deconvolution. The extraction process, designated self-deconvolving data reconstruction algorithm, is applied here to audio-frequency signals showing significant qualitative improvement. Degradation arising from the process of electronic recording and reproduction is significantly reduced.

Acoustics↗

Specification of cross-modal source information in isolated kinematic displays of speech.

Information about the acoustic properties of a talker's voice is available in optical displays of speech, and vice versa, as evidenced by perceivers' ability to match faces and voices based on vocal identity. The present investigation used point-light displays (PLDs) of visual speech and sinewave replicas of auditory speech in a cross-modal matching task to assess perceivers' ability to match faces and voices under conditions when only isolated kinematic information about vocal tract articulation was available. These stimuli were also used in a word recognition experiment under auditory-alone and audiovisual conditions. The results showed that isolated kinematic displays provide enough information to match the source of an utterance across sensory modalities. Furthermore, isolated kinematic displays can be integrated to yield better word recognition performance under audiovisual conditions than under auditory-alone conditions. The results are discussed in terms of their implications for describing the nature of speech information and current theories of speech perception and spoken word recognition.

Acoustic Stimulation↗

Directionality of dog vocalizations.

The directionality patterns of sound emission in domestic dogs were measured in an anechoic environment using a microphone array. Mainly long-distance signals from four dogs were investigated. The radiation pattern of the signals differed clearly from an omnidirectional one with average differences in sound-pressure level between the frontal and rear position of 3-7 dB depending from the individual. Frequency dependence of directionality was shown for the range from 250 to 3200 Hz. The results indicate that when studying acoustic communication in mammals, more attention should be paid to the directionality pattern of sound emission.

Acoustic Stimulation↗