Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sound Spectrography”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

Effects of HearFones on speaking and singing voice quality.

HearFones (HF) have been designed to enhance auditory feedback during phonation. This study investigated the effects of HF (1) on sound perceivable by the subject, (2) on voice quality in reading and singing, and (3) on voice production in speech and singing at the same pitch and sound level. Test 1: Text reading was recorded with two identical microphones in the ears of a subject. One ear was covered with HF, and the other was free. Four subjects attended this test. Tests 2 and 3: A reading sample was recorded from 13 subjects and a song from 12 subjects without and with HF on. Test 4: Six females repeated [pa:p:a] in speaking and singing modes without and with HF on same pitch and sound level. Long-term average spectra were made (Tests 1-3), and formant frequencies, fundamental frequency, and sound level were measured (Tests 2 and 3). Subglottic pressure was estimated from oral pressure in [p], and simultaneously electroglottography (EGG) was registered during voicing on [a:] (Test 4). Voice quality in speech and singing was evaluated by three professional voice trainers (Tests 2-4). HF seemed to enhance sound perceivable at the whole range studied (0-8 kHz), with the greatest enhancement (up to ca 25 dB) being at 1-3 kHz and at 4-7 kHz. The subjects tended to decrease loudness with HF (when sound level was not being monitored). In more than half of the cases, voice quality was evaluated "less strained" and "better controlled" with HF. When pitch and loudness were constant, no clear differences were heard but closed quotient of the EGG signal was higher and the signal more skewed, suggesting a better glottal closure and/or diminished activity of the thyroarytenoid muscle.

Adult↗

The relationship between measured vibrato characteristics and perception in Western operatic singing.

This study examined the association between acoustic and perceptual data related to vibrato in Western operatic singing using recordings of performances by internationally famous opera singers. Three related studies were conducted. Study 1 used commercial recordings of the same five singers and the same cadenza examined by Siegwart and Scherer(1), measured vibrato rate and extent in each singer's performance of the cadenza and tested possible associations between these vibrato attributes and judges' preference for singers. Studies 2 and 3, using recordings of different internationally famous singers and a different cadenza, measured vibrato onset, rate, and extent in each singer's performance of the cadenza, required judges to rank the singers in order of personal preference, to identify the emotion expressed, and to assess the degree of success in communicating emotion achieved by the singer. The findings showed that the perception of the singers' vibrato did not always agree with acoustic measurements. However, a comparison of the acoustic measurements with the preference and emotion judgments suggest that some elements of vibrato may affect listeners' perception of the voice, their preference for a particular singer, and assist the communication of emotion between singer and audience.

Acoustics↗

The perception of two vocal qualities in a synthesized vocal utterance: ring and pressed voice.

Two vocal qualities, ring quality and pressed quality, were analyzed perceptually. Listeners were asked to rate (on a scale from 0 to 10) the "amount of ring" in one listening and the "amount of pressedness" in another listening. The stimulus was the synthesized utterance /ya-ya-ya-ya-ya/. In the continuum representation of ring, the skewing quotient and the cross section of the epilaryngeal tube area were systematically varied, independently and by a covariation rule. In the continuum representation of pressed, the flow amplitude and open quotient were similarly varied. Results indicated that the crossover point between ring and no ring occurred with an epilaryngeal area of around 1.0 cm2, and the crossover point between pressed and not pressed quality occurred at an open quotient of about 0.4. Fundamental frequency also had an effect on the perceptions, with a higher fundamental frequency receiving higher ratings of ring and pressed for otherwise the same parameters. Listeners demonstrated highly variable perceptions in both continua with poor intersubject, intrasubject, and intergroup reliability.

Adolescent↗

Aerodynamic characteristics of laryngectomees breathing quietly and speaking with the electrolarynx.

The primary purpose of this study was to investigate the aerodynamic characteristics of laryngectomees under two conditions: breathing quietly and speaking with electrolarynx. Twenty male adult subjects, 8 normal speakers, and 12 laryngectomees participated the experiment. Airflow, pressure, and speech data were obtained simultaneously. The acceptability of electrolarynx speech under different conditions was also evaluated by 20 listeners (14 men, 6 women). Results indicated a higher peak expiration airflow and pressure among the laryngectomees as compared with the normal during breathing. Three different breathing patterns appeared among the laryngectomees when speaking with the electrolarynx: holding breath, exhaling, and breathing. Four long-time electrolarynx users held breath during speaking. Seven of 12 laryngectomees kept exhaling, whereas only 1 could breathe during speech production. In addition, (1) the acceptability of electrolarynx speech was the highest when speaking breathlessly; (2) no significant difference was found in the acceptability between the patterns of exhaling and breathing smoothly; and (3) the acceptability decreased if breathing quickly during phonation with the electrolarynx. It also suggests that the laryngectomees who can breathe during speaking may be more appropriate to use the new electrolarynx controlling the pitch by expiration pressure.

Adult↗

Effect of nasal decongestion on voice spectrum of a nasal consonant-vowel.

The nasal cavity and its related structures make significant contributions to human phonation, especially the resonance of voice spectra. The voice spectra of the nasal consonant-vowel (CV), [md:], in the subjects with nasal obstruction were obtained and were compared with the spectra of the same CV vocalized by the same subjects after topical nasal decongestion treatment with 1:1000 epinephrine solution. Results revealed that the intensity damping was more marked in the high-frequency area (>1600 Hz) after the nasal decongestion. Moreover, the intensities of the spectral valleys damped more than the spectral peaks, especially the spectral valley of 1000-2700 Hz. Therefore, a more complex spectral pattern was formed by the resultant uneven damping effect after nasal decongestion. The nasal cavity plays an important role in the formation of spectral peaks and valleys, and such engraved voice spectra may also characterize nasal voices like the nasal CV [md:] demonstrated in our study.

Adolescent↗

Acoustic prediction of voice type in women with functional dysphonia.

The categorization of voice into quality type (ie, normal, breathy, hoarse, rough) is often a traditional part of the voice diagnostic. The goal of this study was to assess the contributions of various time and spectral-based acoustic measures to the categorization of voice type for a diverse sample of voices collected from both functionally dysphonic (breathy, hoarse, and rough) (n=83) and normal women (n=51). Before acoustic analyses, 12 judges rated all voice samples for voice quality type. Discriminant analysis, using the modal rating of voice type as the dependent variable, produced a 5-variable model (comprising time and spectral-based measures) that correctly classified voice type with 79.9% accuracy (74.6% classification accuracy on cross-validation). Voice type classification was achieved based on two significant discriminant functions, interpreted as reflecting measures related to "Phonatory Instability" and "F(0) Characteristics." A cepstrum-based measure (CPP/EXP ratio) consistently emerged as a significant factor in predicting voice type; however, variables such as shimmer (RMS dB) and a measure of low- vs. high-frequency spectral energy (the Discrete Fourier Transformation ratio) also added substantially to the accurate profiling and prediction of voice type. The results are interpreted and discussed with respect to the key acoustic characteristics that contributed to the identification of specific voice types, and the value of identifying a subset of time and spectral-based acoustic measures that appear sensitive to a perceptually diverse set of dysphonic voices.

Adult↗

Loud speech in realistic environmental noise: phonetogram data, perceptual voice quality, subjective ratings, and gender differences in healthy speakers.

A new method for cancelling background noise from running speech was used to study voice production during realistic environmental noise exposure. Normal subjects, 12 women and 11 men, read a text in five conditions: quiet, soft continuous noise (75 dBA to 70 dBA), day-care babble (74 dBA), disco (87 dBA), and loud continuous noise (78 dBA to 85 dBA). The noise was presented over loudspeakers and then removed from the recordings in an off-line processing operation. The voice signals were analyzed acoustically with an automatic phonetograph and perceptually by four expert listeners. Subjective data were collected after each vocal loading task. The perceptual parameters press, instability, and roughness increased significantly as an effect of speaking loudly over noise, whereas vocal fry decreased. Having to make oneself heard over noise resulted in higher SPL and F0, as expected, and in higher phonation time. The total reading time was slightly longer in continuous noise than in intermittent noise. The women had 4 dB lower voice SPL overall and increased their phonation time more in noise than did the men. Subjectively, women reported less success making themselves heard and higher effort. The results support the contention that female voices are more vulnerable to vocal loading in background noise.

Acoustics↗

Spectral amplitude measures of adductor spasmodic dysphonic speech.

Spectral amplitude measures are sensitive to varying degrees of vocal fold adduction in normal speakers. This study examined the applicability of harmonic amplitude differences to adductor spasmodic dysphonia (ADSD) in comparison with normal controls. Amplitudes of the first and second harmonics (H1, H2) and of harmonics affiliated with the first, second, and third formants (A1, A2, A3) were obtained from spectra of vowels and /i/ excerpted from connected speech. Results indicated that these measures could be made reliably in ADSD. With the exception of H1(*)-H2(*), harmonic amplitude differences (H1(*)-A1, H1(*)-A2, and H1(*)-A3(*)) exhibited significant negative linear relationships (P < 0.05) with clinical judgments of overall severity. The four harmonic amplitude differences significantly differentiated between pre-BT and post-BT productions (P < 0.05). After treatment, measurements from detected significant differences between ADSD and normal controls (P < 0.05), but measurements from /i/ did not. LTAS analysis of ADSD patients' speech samples proved a good fit with harmonic amplitude difference measures. Harmonic amplitude differences also significantly correlated with perceptual judgments of breathiness and roughness (P < 0.05). These findings demonstrate high clinical applicability for harmonic amplitude differences for characterizing phonation in the speech of persons with ADSD, as well as normal speakers, and they suggest promise for future application to other voice pathologies.

Adult↗

Vocal range and intensity in actors: a studio versus stage comparison.

A voice range profile (VRP) was obtained from each of eight professional actors and compared with two speech range profiles (SRPs). One speech profile was obtained during the dramatic reading of a scene in the laboratory and the other during a performance on stage in a professional theater. The objective was to determine the pitch and loudness ranges used by the actors in speech relative to the VRP. The principal question of interest was whether the actors stayed within the center of the VRP, or whether they tended to drift toward the boundaries of intensity and frequency. A second question was whether the performance within the laboratory accurately reflects that of a stage performance. The results suggest that some subjects tend to exceed the center of the VRP during the stage performance. It is hypothesized that these actors may stress their vocal mechanism during performance and are more likely candidates for vocal injury.

Adult↗

Neck and shoulder muscle activity and thorax movement in singing and speaking tasks with variation in vocal loudness and pitch.

The aim of this study was to examine respiratory phasing and loading levels of sternocleidomastoideus (STM), scalenus (SC), and upper trapezius (TR) muscles in vocalization tasks with variation in vocal loudness and pitch. Eight advanced singing students, aged 22 to 28 years, participated. Surface electromyographic (EMG) activity was recorded from STM, SC, and TR. Thorax movement was detected by two strain gauge sensors placed around the upper (upper TX) and lower (lower TX) thorax. A glissando and simplified singing and speaking tasks were performed. Sustained vowels /a:-i-ae-o:/ were sung in a glissando from lowest to highest pitch (mixed voice/falsetto) back to lowest pitch and in short singing sequences at comfortable, low, and high pitches. The same vowels were spoken softly and loudly for about the same length. The subjects inhaled between the vowels. It was concluded that the inspiratory phased STM and SC muscles produced a counterforce to compression of upper TX at high pitches in glissando. STM and SC were activated to higher levels during phonation than in inhalation. As breathing demands were reduced, STM and SC activity was lowered and the respiratory phasing of peak amplitude changed to inhalation. TR contributed to exhalation in demanding singing with long breathing cycles, but it was less active in singing tasks with short breathing cycles and was essentially inactive in simplified speaking tasks.

Adult↗

The relationship between vocal pitch-matching skills and pitch discrimination skills in untrained accurate and inaccurate singers.

Few studies have compared the relationship between pitch discrimination accuracy and the accuracy of fundamental frequency (F(o)) control. This study investigated the relationship between vocal pitch-matching skills, which is one method of testing F(o) control, and pitch discrimination skills in untrained accurate and inaccurate singers, and the effect of timbre on their pitch discrimination accuracy. Data showed that accurate singers had more precise discrimination and pitch-matching abilities compared with their inaccurate counterparts. Pitch discrimination was differentially affected by the timbre (eg, spectral differences) of comparison tones. In addition, results showed a significant relationship between pitch discrimination abilities and pitch-matching accuracy. The results suggest that accurate F(o) control is at least partially dependent on pitch discrimination abilities, which are important for accurate singing.

Adult↗

Vowel effect on glottal parameters and the magnitude of jaw opening.

This study investigated the relationship among the magnitude of jaw opening, intrinsic fundamental frequency (F0), and glottal parameters in natural speech. Acoustic, jaw opening, and electroglottographic (EGG) signals were simultaneously recorded. The subjects were 10 healthy men with New Zealand English as their native language. Subjects were asked to repeat a standard nonemphasized sentence in which one of the target vowels (/a/, /e/, /i/, /o/, and /u/) was embedded in various contexts. The glottal parameters F0, open quotient (OQ), and speed quotient (SQ) were measured from the EGG signal. Results of a series of one-way repeated-measures analyses of variance (ANOVA) showed a significant vowel effect on the magnitude of jaw opening [F(4, 24) = 25.512, P < .001], F0 [F(4, 28) = 45.415, P < .001] and speed quotient [F(4, 28) = 5.233, P = .003], but not on the open quotient [F(4, 28) = 0.501, P = .735]. The magnitude of jaw opening was found to be inversely related with F0 (r = -0.624, n = 25, P = .0009). These findings showed that the magnitude of jaw opening was related to F0 and that jaw opening might be a control signal for simulation of long-term F0 variation to achieve a higher degree of naturalness in artificial voice.

Adult↗

Resonant voice: spectral and nasendoscopic analysis.

Although resonant voice therapy is a widely used therapeutic approach, little is known about what characterizes resonant voice and how it is physiologically produced. The purpose of this study was to test the hypothesis that resonant voice is produced by narrowing the laryngeal vestibule and is characterized by first formant tuning and more ample harmonics. Videonasendoscopic recordings of the laryngeal vestibule were made during nonresonant and resonant productions of /i/ in six subjects. Spectrums of the two voice types were also obtained. Spectral analysis showed that first formant tuning was exhibited during resonant voice productions and that the degree of harmonic enhancement in the range of 2.0 to 3.5 kHz was related to voice quality: nonresonant voice had the least amount of energy in this range, whereas a resonant-relaxed voice had more energy, and a resonant-bright voice had the greatest amount of energy. Visual-perceptual judgments of the videoendoscopic data indicated that laryngeal vestibule constriction was not consistently associated with resonant voice production.

Adult↗

Perturbation and nonlinear dynamic analyses of voices from patients with unilateral laryngeal paralysis.

This study used perturbation methods (eg, jitter and shimmer) and nonlinear dynamic methods (eg, phase space reconstruction and correlation dimension) to analyze sustained voices generated by normal subjects and patients with unilateral laryngeal paralysis. We found that normal and pathological voices had low-dimensional dynamic characteristics. For nearly periodic voices, jitter and shimmer values of pathological voices from patients with unilateral laryngeal paralysis were significantly different from normal voices. For nearly periodic and aperiodic voices, the correlation dimensions of pathological voices were statistically higher than normal voices. Receiver operating characteristic analysis was used to evaluate the diagnostic performances of jitter, shimmer, and correlation dimension. High sensitivity and specificity of these three acoustic analyses in distinguishing unilateral laryngeal paralysis patients from normal subjects were found. We concluded that combining traditional perturbation analysis and nonlinear dynamic analysis might provide efficient descriptions of pathological voices and represent a valuable tool for clinical diagnosis of laryngeal paralysis.

Adolescent↗

Chaos in voice, from modeling to measurement.

Chaos has been observed in turbulence, chemical reactions, nonlinear circuits, the solar system, biological populations, and seems to be an essential aspect of most physical systems. Chaos may also be central to the interpretation of irregularity in voice disorders. This presentation will summarize the results from a series of our recent studies. These studies have demonstrated the prescence of chaos in computer models of vocal folds, experiments with excised larynges, and human voices. Methods based on nonlinear dynamics can be used to quantify chaos and irregularity in vocal fold vibration. Studies have suggested that disordered voices from laryngeal pathologies such as laryngeal paralysis, vocal polyps, and vocal nodules might exhibit chaotic behaviors. Conventional parameters, such as jitter and shimmer, may be unreliable for analysis of periodic and chaotic voice signals. Nonlinear dynamic methods, however, have differentiated between normal and pathological phonations and can describe the aperiodic or chaotic voice. Chaos theory and nonlinear dynamics can enchance our understanding and therefore our assessment of pathological phonation.

Analysis of Variance↗

M-mode color Doppler ultrasonic imaging of vertical tongue movement during articulatory movement.

To observe and estimate the movement of the tongue, ultrasonic investigation is the most harmless real-time monitoring procedure for analyzing articulatory movements. Color Doppler ultrasonic imaging is special in that it can only sample a moving target, and it can indicate the velocity and direction of the target by color and brightness in real time. This study assessed and demonstrated the validity of M-mode color Doppler ultrasonic imaging to observe the movements of the tongue during syllable repetition tasks performed by normal subjects and dysarthric patients, those affected by amyotrophic lateral sclerosis, cerebellar ataxia, Parkinsonism, and polymyopathy. When the transducer was set below the jaw, upward movement was indicated by a blue signal and downward movement was indicated by a red one on the screen of the ultrasound machine. We also measured the velocity of the tongue by contrast scale classified by 15 degrees. Thus, we could observe vertical tongue movement by a color-coded pattern after quantitative analysis. The Doppler signal patterns of normal subjects were verified by simultaneous video x-ray fluorography recordings. The findings for dysarthric patients corresponded well with previously reported features analyzed by other methods. Therefore, color Doppler ultrasonic imaging of the tongue is a useful procedure to researchers for clinical speech and voice studies.

Dysarthria↗

Acoustic signal typing for evaluation of voice quality in tracheoesophageal speech.

SUMMARY: Because of the aperiodicity of many tracheoesophageal voices, acoustic analysis of the tracheoesophageal voice is less straightforward than that of the normal voice. This study presents the development and testing of an acoustic signal typing system based on visual inspection of a narrow-band spectrogram that can be used by researchers for classification of voice quality in tracheoesophageal speech. In addition to this classification system, a selection of acoustic measures [median fundamental frequency, standard deviation of fundamental frequency, jitter, percentage of voiced (%Voiced), harmonics-to-noise ratio (HNR), glottal-to-noise excitation (GNE) ratio, and band energy difference (BED)] was computed to provide more insight into the acoustic components of tracheoesophageal voice quality. For clinical relevance, relationships between the acoustic signal types and an overall judgment of the voice were investigated as well. Results showed that the four acoustic signal types form a good basis for performing more acoustic analyses and give a good impression of the overall quality of the voice.

Aged↗

The speaker's formant.

The current study concerns speaking voice quality in two groups of professional voice users, teachers (n = 35) and actors (n = 36), representing trained and untrained voices. The voice quality of text reading at two intensity levels was acoustically analyzed. The central concept was the speaker's formant (SPF), related to the perceptual characteristics "better normal voice quality" (BNQ) and "worse normal voice quality" (WNQ). The purpose of the current study was to get closer to the origin of the phenomenon of the SPF, and to discover the differences in spectral and formant characteristics between the two professional groups and the two voice quality groups. The acoustic analyses were long-term average spectrum (LTAS) and spectrographical measurements of formant frequencies. At very high intensities, the spectral slope was rather quandrangular without a clear SPF peak. The trained voices had a higher energy level in the SPF region compared with the untrained, significantly so in loud phonation. The SPF seemed to be related to both sufficiently strong overtones and a glottal setting, allowing for a lowering of F4 and a closeness of F3 and F4. However, the existence of SPF also in LTAS of the WNQ voices implies that more research is warranted concerning the formation of SPF, and concerning the acoustic correlates of the BNQ voices.

Adult↗