Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sound Spectrography”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,117 records · Page 62Linked to original sources

When hearing turns into playing: movement induction by auditory stimuli in pianists.

In this study, pianists were tested for learned associations between actions (movements on the piano) and their perceivable sensory effects (piano tones). Actions were examined that required the playing of two-tone sequences (intervals) in a four-choice paradigm. In Experiment 1, the intervals to be played were denoted by visual note stimuli. Concurrently with these imperative stimuli, task-irrelevant auditory distractor intervals were presented ("potential" action effects, congruent or incongruent). In Experiment 2, imperative stimuli were coloured squares, in order to exclude possible influences of spatial relationships of notes, responses, and auditory stimuli. In both experiments responses in the incongruent conditions were slower than those in the congruent conditions. Also, heard intervals actually "induced" false responses. The reaction time effects were more pronounced in Experiment 2. In nonmusicians (Experiment 3), no evidence for interference could be observed. Thus, our results show that in expert pianists potential action effects are able to induce corresponding actions, which demonstrates the existence of acquired action-effect associations in pianists.

Adult↗

Automated diagnosis of heart disease in patients with heart murmurs: application of a neural network technique.

This study was conducted to test a three-layered artificial neural network analysis of phonocardiogram recordings to diagnose, automatically and objectively, the condition of the heart in patients with heart murmurs. The data were recorded simultaneously in each of 49 patients with a heart murmur through eight microphones attached to the skin surface with adhesive tape, and were analysed by computer. The diagnosis was automated using a three-layered neural network technique. The neural network generated correct answers in over 70% of cases. Furthermore, about 80% of cases of two concurrent diseases were identified correctly. However, ventricular septal defects were incorrectly classified as aortic stenosis or aortic regurgitation, and patent ductus arteriosus was not diagnosed correctly. Accurate diagnoses can frequently be obtained using a neural network, but accuracy can be improved with further data accumulation.

Artificial Intelligence↗

Newborn infants' cry after heel-prick: analysis with sound spectrogram.

The aim of the study was to test the hypothesis that a newborn infant's cry can be used in conjunction with an instrument to measure pain. Crying due to pain was analysed after a heel-prick stimulus. In a prospective, descriptive study, 50 healthy newborn infants were subjected to a heel-prick for phenylketonuria screening. Their cries of pain were recorded and analysed. Duration of the crying sound was analysed and, using a sound spectrogram, the fundamental frequency and the cry melody of the first five cry sounds were analysed. The analysis showed that the crying sound after the painful stimulus of the heel-prick had a significantly higher fundamental frequency and lasted longer at the first than at the fifth cry. The first cry had a more varied crying melody than the fifth. There were large differences between individual cries from a single infant, as well as in the duration of each cry, total crying time, and fundamental frequencies between infants. While the first cry was more like a cry of pain, the fifth cry more resembled crying for reasons other than pain. The results suggest that newborn infants react to pain in a recognizable way. However, other stimuli may cause a similar reaction. Crying can therefore be used to measure pain in newborn infants only when the cause of crying is known.

Blood Specimen Collection↗

Quantitative measurement of speech sound distortions due to inadequate dental mounting.

In this paper, the most adequate quantitative parameters are sought that reflect the dependence of speech sound distortions on the correct position of dental mountings. We suggest the Hurst fractal exponent and some parameters calculated from the Paul wavelet transform of speech sound are useful as quantitative parameters. The investigations are focussed on the alterations of the first /s/ phoneme when the patient produces the Sisyphe sound. The results will be useful to obtain a rigorous, computer assisted method to establish (design) the optimal parameters of prosthetic mounts in order to ensure best speech quality.

Computer Simulation↗

On the differences between conventional and auditory spectrograms of English consonants.

A new tool for speech analysis is presented, operating in real-time and incorporating the analysing power of a contemporary auditory model to produce the familiar display of the speech spectrograph. This "auditory spectrograph" is used to analyse English consonant sounds and the results are compared with conventional wide and narrow band spectrograms. The auditory analyses are found to attach more visual weight to the acoustic cues associated with speech production and perception, and features that are either difficult or impossible to distinguish on conventional spectrograms are clarified.

Female↗

An effect of body massage on voice loudness and phonation frequency in reading.

The effect of massage on voice fundamental frequency (F(0)) and sound pressure level (SPL) was investigated. Subjects were recorded while reading a 3-min passage of prose text. Then, a 30-min session of massage was administered by a trained naprapathy therapist. Sixteen subjects were given the massage, while 15 controls rested, lying down in silence for the same amount of time. The subjects were then recorded reading the same passage again. The F(0), and SPL averages across the whole passage were measured for the pre- and post-treatment recordings. In the post-massage recordings, subjects had lowered their F(0) by 1.1 semitones and their SPL by 1.0 dB, with very high statistical significance. The drop in F(0) was somewhat larger for the males than for the females. The control subjects showed no effect at all.

Adult↗

Long-term average spectrum (LTAS) analysis of sex- and gender-related differences in children's voices.

Long-term average spectrum (LTAS) analysis offers representative information on voice timbre providing spectral information averaged over time. It is particularly useful when persistent spectral features are under investigation. The aim of this study was to compare perceived sex of children to the LTAS analysis of their audio signals. A total of 320 children, aged between 3 and 12 years, were recorded singing a song. In an earlier analysis, the recorded voices were evaluated with respect to perceived and actual sex by experienced listeners. From this group, a subgroup of 59 children (30 boys and 29 girls) was selected. The mean LTAS revealed a peak at 5 kHz for children perceived with confidence as boys, and a flat spectrum at 5 kHz for children perceived confidently as girls (whether male or female in actuality).

Child↗

Spectral distribution of solo voice and accompaniment in pop music.

Singers performing in popular styles of music mostly rely on feedback provided by monitor loudspeakers on the stage. The highest sound level that these loudspeakers can provide without feedback noise is often too low to be heard over the ambient sound level on the stage. Long-term-average spectra of some orchestral accompaniments typically used in pop music are compared with those of classical symphonic orchestras. In loud pop accompaniment the sound level difference between 0.5 and 2.5 kHz is similar to that of a Wagner orchestra. Long-term-average spectra of pop singers' voices showed no signs of a singer's formant but a peak near 3.5 kHz. It is suggested that pop singers' difficulties to hear their own voices may be reduced if the frequency range 3-4 kHz is boosted in the monitor sound.

Acoustics↗

The perception of 'forward' and 'backward placement' of the singing voice.

Singing teachers sometimes characterize voice quality in terms of 'forward' and 'backward placement'. In view of traditional knowledge about voice production, it is hard to explain any possible acoustic or articulatory differences between the voices so 'placed'. We have synthesized a number of three-tone melodic excerpts performed by the singing voice. Formant frequencies, and the level and frequency of the singer's formant were varied across the stimuli. Results of a listening test show that the stimuli which were perceived as 'placed forward', correlated not only with higher frequencies of the first and second formants, but also with the higher frequency and level of the singer's formant.

Female↗

Shared means and meanings in vocal expression of man and macaque.

Vocalisations of six Macaca arctoides that were categorised according to their social context and judgements of naive human listeners as expressions of plea/submission, anger, fear, dominance, contentment and emotional neutrality, were compared with vowel samples extracted from simulations of emotional-motivational connotations by the Finnish name Saara and English name Sarah. The words were spoken by seven Finnish and 13 English women. Humans and monkeys resembled each other in the following respects. 'Neutral', 'pleading' and 'commanding' had a similar F0 level. Loud vocalisations of intense 'anger' and 'fear' had both high F0, and the highest values were encountered for 'fear'. Compared with 'neutral', the audiosignal waveform of 'plea/submission' was more sinusoidal, seen in the spectrum as an attenuation of formants and an emphasis of the fundamental, whereas the signal waveform of 'commanding' was more complex corresponding to an increase in noise and a wider distribution of spectral energy. 'Frightened' samples included rather harmonic segments with emphasis of the fundamental, and 'angry' samples included more noise at the low end of the spectrum and often segments with low-frequency ( < 100 Hz) edged modulation. Sounds resembling soft and noisy 'content' grunts of the monkey do not appear in Finnish or English speech but 'content' utterances were, however, associated with low speech pressure, attenuation of harmonics and increase in noise.

Animals↗

The self-to-other ratio applied as a phonation detector for voice accumulation.

A new method for phonation detection is presented. The method utilises two microphones attached near the subject's ears. Simplified, phonation is assumed to occur when the signals appear mainly in-phase and at equal amplitude. Several signal processing steps are added in order to improve the phonation detection, and finally the original signal is sorted in separate channels corresponding to the phonated and non-phonated instances. The method is tested in a laboratory setting to demonstrate the need for some of the stages of the signal processing and to examine the processing speed. The resulting sound file allows for measurement of phonation time, speaking time and fundamental frequency of the subject and sound pressure level of the subject's voice and the environmental sounds separately. The present implementation gives great freedom for adjustment of analysis parameters, since the microphone signals are recorded on DAT tape and the processing is performed off-line on a PC. In future versions, a voice accumulator based on this principle could be designed in order to shorten analysis time and thus make the method more appropriate for clinical use.

Equipment Design↗

Source and filter adjustments affecting the perception of the vocal qualities twang and yawn.

Two vocal qualities, twang and yawn, were synthesized and rated perceptually. The stimuli consisted of synthesized vocal productions of a sentence-length utterance 'ya ya ya ya ya,' which had speech-like intonation. In a continuum transformation from normal to twang, the area in the pharynx was gradually decreased, along with vocal tract shortening and a decreased open quotient in the glottal airflow. In a continuum transformation toward yawn, the area in the pharynx was gradually increased, along with vocal tract lengthening and an increased open quotient. The normal (untransformed) vocal tract area was pre-determined by earlier studies involving MRI scans of a human subject's vocal tract. Listeners were asked to rate (on a scale from 1-10) the 'amount of twang' in one listening session and the 'amount of yawn' in another listening session. Overall, the perception of twang increased directly with pharyngeal area narrowing, vocal tract shortening, and decreased open quotient. The perception of yawn increased with pharyngeal area widening, vocal tract lengthening, and increased open quotient. Adjustments of one parameter alone yielded less significant perceptual changes than the above combinations, with open quotient showing the greatest effect in isolation. Listeners demonstrated variable perceptions in both continua with poor inter-subject, intra-subject, and inter-group reliability.

Adult↗

Measurement of vocal doses in speech: experimental procedure and signal processing.

An experimental method for quantifying the amount of voicing over time is described in a tutorial manner. A new procedure for obtaining calibrated sound pressure levels (SPL) of speech from a head-mounted microphone is offered. An algorithm for voicing detection (kv) and fundamental frequency (F0) extraction from an electroglottographic signal is described. The extracted values of SPL, F0, and kv are used to derive five vocal doses: the time dose (total voicing time), the cycle dose (total number of vocal fold oscillatory cycles), the distance dose (total distance travelled by the vocal folds in an oscillatory path), the energy dissipation dose (total amount of heat energy dissipated in the vocal folds) and the radiated energy dose (total acoustic energy radiated from the mouth). The doses measure the vocal load and can be used for studying the effects of vocal fold tissue exposure to vibration.

Calibration↗

WinSingad: a real-time display for the singing studio.

This paper describes the nature and implementation of a specially-designed integrated real-time display that is undergoing evaluation as part of a recently funded innovative pilot project to investigate the relative usefulness of computer displays in the singing studio. Following previous work that suggests that simple displays of a small number of analysis parameters are generally likely to be the most effective, the system makes available a range of complementary analyses that are plotted against time. These relate to: fundamental frequency, spectrum, spectral ratio, and vocal tract area. These can be viewed singly, multiply or in combination using a panel based design within the PC Windows environment, known as WinSingad. The algorithms used are described and the displays themselves are illustrated with results gained from the pilot phase of the research to indicate their potential usefulness.

Adult↗

The impact of 'open throat' technique on vibrato rate, extent and onset in classical singing.

Mitchell, Kenny et al. (2003) identified 'open throat' as integral to the production of an even and consistent sound in classical singing. In this study, we compared vibrato rate, extent and onset of six advanced singing students under three conditions: 'optimal' (O), representing maximal open throat; 'sub-optimal' (SO), using reduced open throat; and loud sub-optimal (LSO), using reduced open throat but controlling for the effect of loudness. Fifteen expert judges correctly identified the sound produced when singers used open throat with 85% accuracy. Having verified the technique perceptually, we used a series of univariate repeated measures ANOVAs with planned orthogonal contrasts to test the hypotheses that frequency modulations associated with vibrato rate, extent and onset would vary outside acceptable or desirable parameters for SO and LSO. Hypotheses were confirmed for vibrato extent and onset but not for rate. There were no significant differences between SO and LSO on any of the vibrato parameters. As vibrato is considered a key indicator of good singing, these findings suggest that open throat is important to the production of a good sound in classical singing.

Adult↗

Effect on LTAS of vocal loudness variation.

Long-term-average spectrum (LTAS) is an efficient method for voice analysis, revealing both voice source and formant characteristics. However, the LTAS contour is non-uniformly affected by vocal loudness. This variation was analyzed in 15 male and 16 female untrained voices reading a text 7 times at different degrees of vocal loudness, mean change in overall equivalent sound level (Leq) amounting to 27.9 dB and 28.4 dB for the female and male subjects. For all frequency values up to 4 kHz, spectrum level was strongly and linearly correlated with Leq for each subject. The gain factor, that is to say, the rate of level increase, varied with frequency, from about 0.5 at low frequencies to about 1.5 in the frequency range 1.5-3 kHz. Using the gain factors for a subject, LTAS contours could be predicted at any Leq within the measured range, with an average accuracy of 2-3 dB below 4 kHz. Mean LTAS calculated for an Leq of 70 dB for each subject showed considerable individual variation for both males and females, SD of the level varying between 7 dB and 4 dB depending on frequency. On the other hand, the results also suggest that meaningful comparisons of LTAS, recorded for example before and after voice therapy, can be made, provided that the documentation includes a set of recordings at different loudness levels from one recording session.

Adult↗

The effects of open throat technique on long term average spectra (LTAS) of female classical voices.

In the third of a series of studies on open throat technique, we compared long term average spectra (LTAS) of six advanced singing students under three conditions: 'optimal' (O), representing maximal open throat, 'sub-optimal' (SO), using reduced open throat, and loud sub-optimal (LSO) to control for the effect of loudness. Using a series of univariate repeated measures ANOVAs with planned orthogonal contrasts, we tested the hypotheses that sound pressure level (SPL) and the ratio of spectral energy in peaks and areas between 0-2 kHz and 2-4 kHz would be reduced in SO and LSO compared to O. There were significant differences between SO and LSO but hypotheses were not confirmed for O. These findings do not accord with differences in vibrato extent and onset between O and SO/LSO (Mitchell and Kenny, in press). These results suggest that while LTAS provides information on energy distribution, measuring spectral energy areas appears to be the most sensitive measure of energy distribution between conditions. Plotting the differences between O and SO/LSO pairs of LTAS clearly indicates the areas of spectral change. The findings from this study also indicate that LTAS are not sufficiently sensitive to measure vocal timbre as they were not consistent with perceptual or other acoustic studies of the same samples.

Acoustics↗

Vocal fold vibration and voice source aperiodicity in 'dist' tones: a study of a timbral ornament in rock singing.

The acoustic characteristics of so-called 'dist' tones, commonly used in singing rock music, are analyzed in a case study. In an initial experiment a professional rock singer produced examples of 'dist' tones. The tones were found to contain aperiodicity, SPL at 0.3 m varied between 90 and 96 dB, and subglottal pressure varied in the range of 20-43 cm H2O, a doubling yielding, on average, an SPL increase of 2.3 dB. In a second experiment, the associated vocal fold vibration patterns were recorded by digital high-speed imaging of the same singer. Inverse filtering of the simultaneously recorded audio signal showed that the aperiodicity was caused by a low frequency modulation of the flow glottogram pulse amplitude. This modulation was produced by an aperiodic or periodic vibration of the supraglottic mucosa. This vibration reduced the pulse amplitude by obstructing the airway for some of the pulses produced by the apparently periodically vibrating vocal folds. The supraglottic mucosa vibration can be assumed to be driven by the high airflow produced by the elevated subglottal pressure.

Air Pressure↗