Search PubMedSearch

Biomedical subjects

E Terhardt

Publications and source records attributed to E Terhardt.

3 recordsLinked to original sources

Calculating virtual pitch.

A procedure for the schematic and automatic extraction of 'fundamental pitch' from complex tonal signals, such as voiced speech and music, has been developed. While the auditively relevant 'fundamental' of a complex signal cannot be defined in purely mathematical terms, an existent model of virtual-pitch perception turns out to provide a suitable basis. The procedure comprises the formation of determinant spectral pitches (or 'fundamental frequency') from those spectral pitches. The latter deduction is accomplished by a principle of subharmonic matching, for whose realization a simple, universal and efficient algorithm was found. While the calculation may be confined to the determination of 'nominal' virtual pitch, certain typical auditory phenomena, such as the influence of SPL, partial masking and interval stretch, may be accounted for as well, in which case 'true' virtual pitch is obtained. The procedure operates on the frequencies and amplitudes of the signal's spectral components, is suitable for implementation on readily available programmable calculators and other arithmetic computers, and may be used in real-time 'fundamental-pitch' extraction as well. The procedure's performance and its applicability to the research and engineering of auditory communication are illustrated by some examples.

Humans

Automatic speech recognition using psychoacoustic models.

An approach to automatic speech recognition is described, which, in a straightforward way, follows the concept of (1) preprocessing in terms of auditory parameters and (2) subsequent classification and recognition. The preprocessing system has been realized in analog hardware, while recognition is carried out on a digital computer. In the preprocessing system, the essential psychoacoustic principles of the perception of loudness, pitch, roughness, and subjective duration are implemented with some approximation. The system essentially consists of 24 bandpass filters, nonlinear transformation of each filter output into specific loudness and specific roughness, and final transformation of these parameters into total loudness, total roughness, and three spectral momenta. As a means to further reduce the information flow, continuous selection of dominant parameters is also considered on the basis of psychoacoustic data. The subsequent recognition process is mainly characterized by (1) discrimination between speech and silent periods, (2) detection of syllable peaks and classification of syllable nuclei, and (3) assumption of syllable boundaries and classification of consonant clusters. Though the entire system as yet is far from being complete and perfect, the present results indicate that the concept provides a systematic and promising way towards automatic recognition of continuous speech.

Computers