Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sound Spectrography”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

The acoustic characteristics of professional opera singers performing in chorus versus solo mode.

In this study, members of a professional opera chorus were recorded using close microphones, while singing in both choral and solo modes. The analysis included computation of long-term average spectra (LTAS) for the two song sections performed and calculation of singing power ratio (SPR) and energy ratio (ER), which provide an indication of the relative energy in the singer's formant region. Vibrato rate and extent were determined from two matched vowels, and SPR and ER were calculated for these vowels. Subjects sung with equal or more power in the singer's formant region in choral versus solo mode in the context of the piece as a whole and in individual vowels. There was no difference in vibrato rate and extent between the two modes. Singing in choral mode, therefore, required the ability to use a similar vocal timbre to that required for solo opera singing.

Cooperative Behavior↗

Vocal expression of emotions in normally hearing and hearing-impaired infants.

The vocalizations of seven normally hearing (NH) and seven severely hearing-impaired (HI) infants were compared to find out the influence of auditory feedback on preverbal utterances. It was tested whether there are general differences in vocalizations between NH and HI infants, and whether specific emotional states affect the vocal production of NH and HI infants in the same way. First, the acoustic structure of the three most common vocal types was analyzed; second, the composition of vocal sequences was examined. Vocal sequence composition turned out to be more affected by hearing impairment than the acoustic structure of single vocalizations. This result indicates that the acoustic structure of preverbal vocalizations is to a great extent predetermined, whereas the composition of vocal sequences is influenced by auditory input.

Affect↗

Are real-time displays of benefit in the singing studio? An exploratory study.

This article reports on an exploratory research project to evaluate the usefulness or otherwise of real-time visual feedback in the singing studio. The primary purpose of the work was not to optimize the technology for this application, but to work alongside teachers and students to study the impact of real-time visual feedback technology use on the students' learning experiences. An action research methodology was used to explore the benefit of real-time displays over an extended period. The experimental phase of the work was guided by a Liaison Panel of teachers and academics in the areas of singing, pedagogy, voice science, speech therapy, and linguistic science. Qualitative data were collected from eight students working with two professional singing teachers. The teachers and students acted as co-researchers under the action research paradigm. Teachers and students alike kept journals of their teaching and learning experiences. Singing lessons were observed regularly by the research team, coded for teacher and student behaviors, and all co-researchers were interviewed at the mid- and endpoint of the project. The use of technology had a positive impact on the learning process, and this is evidenced through case study data.

Equipment Design↗

Intonation drift in a capella soprano, alto, tenor, bass quartet singing with key modulation.

OBJECTIVES/HYPOTHESIS: When a soprano, alto, tenor, bass (SATB) quartet sings unaccompanied, or a capella, the members of the group will tend to make use of non-equal-tempered intonation to govern their tuning. If the music they are performing visits different keys and they do maintain non-equal-tempered tuning, then the pitch center will have to shift from its starting point, which is a necessary consequence of the physics behind the use of a non-equal-tempered tuning system. The implication of this shift for tuning in a capella singing is that it is not possible both to maintain accurate non-equal-tempered tuning and to stay in pitch throughout music that modulates in key. METHODS: To test this notion, a set of four-part exercises were written by the author that visit several different key chords in sequential progression. In each case, the starting and finishing chords were either identical or exactly an octave apart. Mean fundamental frequency values for each note were measured using four electrolaryngographs (one per singer), and the f0 data were normalized and plotted with respect to equal-tempered tuning to enable any overall tuning shift to be observed. RESULTS: The results indicate that singers do (1) tend to non-equal-tempered tuning and (b) do consequentially shift their intonation with modulation. CONCLUSIONS: These data indicate that pitch drift is potentially a necessary part of staying in-tune. Further work is required to identify items in the choral repertoire for which this effect is likely and then to inform the choral conducting and singing communities appropriately.

Cooperative Behavior↗

Frequency and voice: perspectives in the time domain.

SUMMARY: Frequency variation is one of the most primitive features of voice production, endowing language and communication with richness and efficiency and enhancing enjoyment of the voice arts. In the first of two tutorial articles, the subject of frequency is examined formally, beginning in the time domain. A companion article explores the topic of frequency and voice from the frequency domain perspective. Frequency is a well-defined quantity of the sinusoidal function and of periodic functions of time. However, voice is inherently nonstationary, even over short time segments, to degrees that range from minor (stable vowels of a healthy voice) to major (singing voice and voiced consonants). For signals that are not periodic, the notion of frequency is ambiguous and often altogether unclear, which has led to a multitude of frequency-measurement techniques and discrepancy of measures. This article identifies the source of these discrepancies for a variety of time-domain techniques that are examined in the absence of noise. In the time domain, the subject of frequency is inherently coupled to the topic of signal modeling, which is explored in some detail. Sinusoidal models having time-varying phase are examined with the objective of achieving a frequency description of voice that is both continuous and instantaneous. The analytic signal method of mathematical physics is discussed and applied to the technology of empirical mode decomposition to demonstrate that the frequencies of voice may be comprehensively examined from the time domain point of view.

Humans↗

When does a sung tone start?

Although the consonant is mostly considered as the start of a syllable in phonetics and orthography, musicians generally agree that the vowel onset in singing should be synchronized with the beat. As a test of this assumption, the current investigation analyzes the time interval between vowel onsets and piano accompaniment onsets in a set of songs performed by international vocal artists and published on commercial CD recordings. The results show that, most commonly, the accompanists synchronized their tones with the singers' vowel onsets. Nevertheless, examples of lead and lag were found, probably made for expressive purposes. The lead and lag varied greatly between songs, being smallest in a song performed in a fast tempo and longest in a song performed in a slow tempo.

Humans↗

Reliability of speaking and maximum voice range measures in screening for dysphonia.

Speech range profile (SRP) is a graphical display of frequency-intensity occurring interactions during functional speech activity. Few studies have suggested the potential clinical applications of SRP. However, these studies are limited to qualitative case comparisons and vocally healthy participants. The present study aimed to examine the effects of voice disorders on speaking and maximum voice ranges in a group of vocally untrained women. It also aimed to examine whether voice limit measures derived from SRP were as sensitive as those derived from voice range profile (VRP) in distinguishing dysphonic from healthy voices. Ninety dysphonic women with laryngeal pathologies and 35 women with normal voices, who served as controls, participated in this study. Each subject recorded a VRP for her physiological vocal limits. In addition, each subject read aloud the "North Wind and the Sun" passage to record SRP. All the recordings were captured and analyzed by Soundswell's computerized real-time phonetogram Phog 1.0 (Hitech Development AB, Täby, Sweden). The SRPs and the VRPs were compared between the two groups of subjects. Univariate analysis results demonstrated that individual SRP measures were less sensitive than the corresponding VRP measures in discriminating dysphonic from normal voices. However, stepwise logistic regression analyses revealed that the combination of only two SRP measures was almost as effective as a combination of three VRP measures in predicting the presence of dysphonia (overall prediction accuracy: 93.6% for SRP vs 96.0% for VRP). These results suggest that in a busy clinic where quick voice screening results are desirable, SRP can be an acceptable alternate procedure to VRP.

Adult↗

Effect of deep brain stimulation on different speech subsystems in patients with multiple sclerosis.

The effect of deep brain stimulation on articulation and phonation subsystems in seven patients with multiple sclerosis (MS) was examined. Production parameters in fast syllable-repetitions were defined and measured, and the phonation quality during vowel productions was analyzed. Speech material was recorded for patients (with and without stimulation) and for a group of healthy control speakers. With stimulation, the precision of glottal and supraglottal articulatory gestures is reduced, whereas phonation has a greater tendency to be hyperfunctional in comparison with the healthy control data. Different effects on the two speech subsystems are induced by electrical stimulation of the thalamus in patients with MS.

Adult↗

Effects of professional singing education on vocal vibrato--a longitudinal study.

Vocal vibrato is regarded as one of the essential characteristics of voice quality in classical singing. Professional singers seem to develop vibrato automatically, without actively striving to acquire it. In this longitudinal investigation, the vocal vibrato of 22 singing students was examined at the beginning of and after 3 years of professional singing education. Subjects sang an ascending-descending triad pattern in slow tempo on vowel [a:] at a comfortable pitch level twice at soft (piano) and twice at medium (mezzoforte) loudness. The top note of the triad pattern was sustained for approximately 5s. The mean and the standard deviation (SD) of the vibrato rate were measured for this note. Results revealed that after 3 years of training, voices with vibrato slower than 5.2 Hz were found to have a faster vibrato, and voices with vibrato faster than 5.8 Hz were found to have a slower vibrato. Standard deviation of vibrato rate was higher in soft than in medium loudness, particularly before the education. Also high values of SD of vibrato rate, exceeding 0.65 Hz, had decreased after the education. These findings confirm that vibrato characteristics can be affected by singing education.

Adult↗

Assessment of the formant frequencies in normal and laryngectomized individuals using linear predictive coding.

The objective of this study was to assess the difference in voice quality as defined by acoustical analysis using sustained vowel in laryngectomized patients in comparison with normal volunteers. This was designed as a retrospective single center cohort study. An adult tertiary referral unit formed the setting of this study. Fifty patients (40 males) who underwent total laryngectomy and 31 normal volunteers (18 male) participated. Group comparisons with the first three formant frequencies (F1, F2, and F3) using linear predictive coding (LPC) (Laryngograph Ltd, London, UK) was performed. The existence of any significant difference of F1, F2, and F3 between the two groups using the sustained vowel /i/ and the effects of other factors namely, tumor stage (T), chemoradiotherapy, pharyngectomy, cricothyroid myotomy, closure of pharyngoesophageal segment, and postoperative complication were analyzed. Formant frequencies F1, F2, and F3 were significantly different in male laryngectomees compared to controls: F1 (P<0.001, Mann-Whitney U test), F2 (P<0.001, Student's t test), and F3 (P=0.008, Student's t test). There was no significant difference between females in both groups for all three formant frequencies. Chemoradiotherapy and postoperative complications (pharyngocutaneous fistula) caused a significantly lower formant F1 in men, but showed little effect in F2 and F3. Laryngectomized males produced significantly higher formant frequencies, F1, F2, and F3, compared to normal volunteers, and this is consistent with literature. Chemoradiotherapy and postoperative complications significantly influenced the formant scores in the laryngectomee population. This study shows that robust and reliable data could be obtained using electroglottography and LPC in normal volunteers and laryngectomees using a sustained vowel.

Female↗

Changes in objective acoustic measurements and subjective voice complaints in call center customer-service advisors during one working day.

SUMMARY: The aim of this study was to investigate how different acoustic parameters, extracted both from speech pressure waveforms and glottal flows, can be used in measuring vocal loading in modern working environments and how these parameters reflect the possible changes in the vocal function during a working day. In addition, correlations between objective acoustic parameters and subjective voice symptoms were addressed. The subjects were 24 female and 8 male customer-service advisors, who mainly use telephone during their working hours. Speech samples were recorded from continuous speech four times during a working day and voice symptom questionnaires were completed simultaneously. Among the various objective parameters, only F0 resulted in a statistically significant increase for both genders. No correlations between the changes in objective and subjective parameters appeared. However, the results encourage researchers within the field of occupational voice use to apply versatile measurement techniques in studying occupational voice loading.

Adult↗

Acoustic analysis of consonants in whispered speech.

An acoustic analysis of whispered consonants in comparison to normally phonated consonants was conducted in time and intensity domains. Consonant duration and average root mean square intensity were measured for six speakers in both articulation modes. Each of 25 Serbian consonants (C) was sited between the vowel /a/ forming a syllable of /aCa/ type. Such a syllable was placed in initial, medial, and final position in the carrier sentence. Results showed that whispered consonants have a prolonged duration of about 10% on average (statistically significant, ANOVA test), and that the unvoiced consonants have a smaller time dimension extension (5.8%) than voiced ones (15.3%). Examination at subphonemic level showed that there is no difference in voice-onset-time and affrication duration in unvoiced plosives and affricates, in both whispered and phonated mode of articulation, but the difference is significant for voiced ones. Analysis of consonant duration versus place of articulation showed that palatal place is most sensitive in the process of whispering. In all experiments, the results are very consistent with respect to the subjects and test material (Pearson's correlation was between 0.6 and 0.9). In intensity domain, all unvoiced consonants in whispered mode of articulation have almost unchanged intensity in comparison to phonated mode (the difference is maximum 3.5 dB). On the contrary, voiced consonants in the whispered mode were reduced in intensity by as much as 25 dB, as nasals and semivowels. Average intensity of whispered consonants is lowered by 12d B in comparison to phonated ones, and does not depend on syllabic position inside the sentences.

Adult↗

Acoustic and perceptual analyses of Brazilian male actors' and nonactors' voices: long-term average spectrum and the "actor's formant".

SUMMARY: This study investigates the possible differences between actors' and nonactors' vocal projection strategies using acoustic and perceptual analyses. A total of 11 male actors and 10 male nonactors volunteered as subjects, reading an extended text sample in habitual, moderate, and loud levels. The samples were analyzed for sound pressure level (SPL), alpha ratio (difference between the average SPL of the 1-5kHz region and the average SPL of the 50Hz-1kHz region), fundamental frequency (F0), and long-term average spectrum (LTAS). Through LTAS, the mean frequency of the first formant (F1) range, the mean frequency of the "actor's formant," the level differences between the F1 frequency region and the F0 region (L1-L0), and the level differences between the strongest peak at 0-1kHz and that at 3-4kHz were measured. Eight voice specialists evaluated perceptually the degree of projection, loudness, and tension in the samples. The actors had a greater alpha ratio, stronger level of the "actor's formant" range, and a higher degree of perceived projection and loudness in all loudness levels. SPL, however, did not differ significantly between the actors and nonactors, and no differences were found in the mean formant frequencies ranges. The alpha ratio and the relative level of the "actor's formant" range seemed to be related to the degree of perceived loudness. From the physiological point of view, a more favorable glottal setting, providing a higher glottal closing speed, may be characteristic of these actors' projected voices. So, the projected voices, in this group of actors, were more related to the glottic source than to the resonance of the vocal tract.

Adult↗

Source-filter comparison of measurements of fundamental frequency perturbation and amplitude perturbation for synthesized voice signals.

SUMMARY: An investigation of the effect of glottal source aperiodicities (jitter, shimmer, and aspiration noise) on the estimation of fundamental frequency (f0) perturbation and amplitude perturbation, of synthesized, glottal source and voiced speech waveforms, is considered. Firstly, 4, cycle-event f0 estimators are examined: (1) waveform matching of the low-pass filtered waveform, (2) positive peaks (PPs) from the speech waveform, (3) PPs from the low-pass filtered waveform, and (4) positive zero crossings from the low-pass filtered waveform. The analysis shows that f0 perturbation measures taken from the low-pass filtered waveform are affected by both amplitude perturbation and random glottal noise, whereas, f0 perturbation measures taken from the PPs of the original waveform are affected by noise but not by amplitude perturbation. It is shown for the low-pass filter methods that the effects of amplitude perturbation and noise lead to increased errors in the measurement of f0 perturbation for the synthesized speech waveforms when compared with the synthesized glottal waveforms. Shimmer of the synthesized speech waveform is approximately equal to shimmer of the synthesized glottal source. However, noise and jitter affect measures of amplitude perturbation. The estimation of f0 perturbation from the synthesized speech waveform is shown to be nonlinearly related to f0 perturbation estimation from the synthesized glottal waveform as a consequence of the filtering action of the vocal tract. Low-pass filtering the voiced speech waveform is shown to provide a partial solution to this problem.

Communication Devices for People with Disabilities↗

Voice of postradiotherapy nasopharyngeal carcinoma patients: evidence of vocal tract effect.

This study was aimed at identifying acoustic and physiological measures useful for monitoring voice changes in postnasopharyngeal patients with nonlaryngeal malignancies, and providing evidences of vocal tract effect on voice through comparisons between individuals with and without intact vocal tract. Simultaneous acoustic-electroglottographic signals recorded during phonation of vowels /i/ and /a/ sustained at habitual, high, and low pitch levels were compared among 10 postradiotherapy patients with nasopharyngeal carcinoma (NPC), 10 voice patients (VPs) with intact vocal tract, and 10 healthy individuals with normal voice (NORM). Results from a series of discriminant analyses revealed that the NPC group generally exhibited lower signal-to-noise (SNR) and open quotient (OQ) and higher Formant 1 frequency (F(1)) and speed quotient (SQ) than the NORM group. Unlike both VP and NORM groups, the NPC group failed to show a pitch effect on all voice measures, including OQ, SQ, percent jitter, percent shimmer, and SNR, suggesting an effect of radiotherapy and/or vocal tract on laryngeal behaviors. For the vowel /i/, on the other hand, only the NPC and NORM groups showed a pattern of pitch-dependent F(1) raising, a reflection of increased pharyngeal narrowing. These findings suggested that the pitch effect on laryngeal behaviors differed not only between individuals with intact vocal tract and those without but also between those with structural and dynamic changes of vocal tract.

Adult↗

Acoustic measures and self-reports of vocal fatigue by female teachers.

This study investigated the relation of symptoms of vocal fatigue to acoustic variables reflecting type of voice production and the effects of vocal loading. Seventy-nine female primary school teachers volunteered as subjects. Before and after a working day, (1) a 1-minute text reading sample was recorded at habitual loudness and loudly (as in large classroom), (2) a prolonged phonation on [a:] was recorded at habitual speaking pitch and loudness, and (3) a questionnaire about voice quality, ease, or difficulty of phonation and tiredness of throat was completed. The samples were analyzed for average fundamental frequency (F0), sound pressure level (SPL), and phonation type reflecting alpha ratio (SPL [1-5 kHz]-SPL [50 Hz-1 kHz]). The vowel samples were additionally analyzed for perturbation (jitter and shimmer). After a working day, F0, SPL, and alpha ratio were higher, jitter and shimmer values were lower, and more tiredness of throat was reported. The average levels of the acoustic parameters did not correlate with the symptoms. Increase in jitter and mean F0 in loud reading correlated with tiredness of throat. The results seem to suggest that, at least among experienced vocal professionals, voice production type had little relevance from the point of view of vocal fatigue reported. Differences in the acoustic parameters after a vocally loading working day mainly seem to reflect increased muscle activity as a consequence of vocal loading.

Adult↗

Dissimilarity and the classification of male singing voices.

Traditionally, timbre has been defined as that perceptual attribute that differentiates two sounds when pitch and loudness are equal, and thus is a measure of dissimilarity. By such a definition, each voice possesses a set of timbres, and the ability to identify any voice or voice category across different pitch-loudness-vowel combinations must be due to an ability to "link" these timbres by abstracting the "timbre transformation," the manner in which timbre subtly changes across pitch and loudness for a specific voice or voice category. Using stimuli produced across the singing range by singers from different voice categories, this study sought to examine how timbre and pitch interact in the perception of dissimilarity in male singing voices. This study also investigated whether or not listener experience affects the perception of timbre as a function of pitch. The resulting multidimensional scaling (MDS) representations showed that for all stimuli and listeners, dimension 1 correlated with pitch, while dimension 2 correlated with spectral centroid and separated vocal stimuli into the categories baritone and tenor. Dimension 3 appeared highly idiosyncratic depending on the nature of the stimuli and on the experience of the listener. Inexperienced listeners appeared to rely more heavily on pitch in making dissimilarity judgments than did experienced listeners. The resulting MDS representations of dissimilarity across pitch provide a glimpse of the timbre transformation of voice categories across pitch.

Humans↗

The influence of underwater data transmission sounds on the displacement behaviour of captive harbour seals (Phoca vitulina).

To prevent grounding of ships and collisions between ships in shallow coastal waters, an underwater data collection and communication network (ACME) using underwater sounds to encode and transmit data is currently under development. Marine mammals might be affected by ACME sounds since they may use sound of a similar frequency (around 12 kHz) for communication, orientation, and prey location. If marine mammals tend to avoid the vicinity of the acoustic transmitters, they may be kept away from ecologically important areas by ACME sounds. One marine mammal species that may be affected in the North Sea is the harbour seal (Phoca vitulina). No information is available on the effects of ACME-like sounds on harbour seals, so this study was carried out as part of an environmental impact assessment program. Nine captive harbour seals were subjected to four sound types, three of which may be used in the underwater acoustic data communication network. The effect of each sound was judged by comparing the animals' location in a pool during test periods to that during baseline periods, during which no sound was produced. Each of the four sounds could be made into a deterrent by increasing its amplitude. The seals reacted by swimming away from the sound source. The sound pressure level (SPL) at the acoustic discomfort threshold was established for each of the four sounds. The acoustic discomfort threshold is defined as the boundary between the areas that the animals generally occupied during the transmission of the sounds and the areas that they generally did not enter during transmission. The SPLs at the acoustic discomfort thresholds were similar for each of the sounds (107 dB re 1 microPa). Based on this discomfort threshold SPL, discomfort zones at sea for several source levels (130-180 dB re 1 microPa) of the sounds were calculated, using a guideline sound propagation model for shallow water. The discomfort zone is defined as the area around a sound source that harbour seals are expected to avoid. The definition of the discomfort zone is based on behavioural discomfort, and does not necessarily coincide with the physical discomfort zone. Based on these results, source levels can be selected that have an acceptable effect on harbour seals in particular areas. The discomfort zone of a communication sound depends on the sound, the source level, and the propagation characteristics of the area in which the sound system is operational. The source level of the communication system should be adapted to each area (taking into account the width of a sea arm, the local sound propagation, and the importance of an area to the affected species). The discomfort zone should not coincide with ecologically important areas (for instance resting, breeding, suckling, and feeding areas), or routes between these areas.

Acoustics↗