|
Author: Gerald Schuller Co-Author: Jürgen Herre
The goal of a high compression ratio in perceptual audio coding has historically led to the use of transforms with a large block size or filter banks with many bands (i.e. high frequency resolution). Such spectral decompositions are suitable to obtain high coding gains for the mostly stationary parts in music signals. On the other hand, due to the so-called uncertainty principle, a high frequency resolution implies a low temporal resolution of the time/frequency representation. Thus, a high frequency resolution results in a poor control over the temporal shape of the quantization noise in the decoded audio signal, which may lead to coding artifacts for less stationary audio material. This can be perceived, for instance, as reverberation or echoiness in speech signals or pre-echoes at the "attack" portions of transient signals, such as castanets. More background information on this effect can be found in the section on the Pre-Echo phenomenon. To obtain a better control over the temporal shaping of the quantization noise, a number of techniques were proposed over time:
Sound examples The signals were generated using the popular Modified Discrete Cosine Transform (MDCT) [5], which is used in many of today's coding schemes and a sine window. There are two parameters which were varied across the different signal versions:
As the size of the filter bank window increases, a better frequency resolution is achieved at the expense of a decreased temporal resolution. The processed signals contain a controlled level of noise over frequency, which is indicated relative to an (arbitrarily chosen) reference level. Processed signals with lower noise levels may be used to subsequently train critical listening with more subtle versions of the artifacts.
The following effects can be noticed when listening to the sound excerpts: The temporal "smearing" of the distortion introduces a reverberant quality into the speech signal which increases significantly with the length of the filter bank window. For large window sizes, the effect is even audible at rather small distortion levels. Accordingly, proper use of additional measures for preventing temporal unmasking is of high importance for audio coders with a high frequency resolution / number of filter bank channels. To demonstrate the character of this coding artifact with
a real coder, the number of subbands in an audio coder was artificially
fixed to 1024, to avoid switching to its 128 band mode. This increases
the coding artifacts to make them more audible. The test signal again
consists of German male speech, a typical test signal were artifacts are
easily produced and detected. The signal has been coded at a sampling
rate of 32 kHz and a bit-rate of 64 kb/s (for the stereo file). Again,
it is recommended that these samples are heard over headphones, otherwise
the real room reverberation might mask the artifacts: References: [1] B. Edler: "Codierung von Audiosignalen mit überlappender Transformation und adaptiven Fensterfunktionen", Frequenz, Vol. 43, pp. 252-256, 1989 [2] J. Herre, J. D. Johnston: "Enhancing the Performance of Perceptual Audio Coders by Using Temporal Noise Shaping (TNS)", 101st AES Convention, Los Angeles 1996, Preprint 4384 [3] T. Vaupel: "Ein Beitrag zur Transformationscodierung von Audiosignalen unter Verwendung der Methode der 'Time Domain Aliasing Cancellation (TDAC)' und einer Signalkompandierung im Zeitbereich", PhD Thesis, Universität-Gesamthochschule Duisburg, Germany, 1991 [4] B. Edler, C. Faller, G. Schuller: "Perceptual Audio Coding Using a Time-Varying Linear Pre- and Post-Filter", 109th AES Convention, Los Angeles 2000, Preprint 5274 [5] J. Princen, A. Johnson, A. Bradley: "Subband/Transform
Coding Using Filter Bank Designs Based on Time Domain Aliasing Cancellation",
IEEE ICASSP 1987, pp. 2161 - 2164 |
||||||||||||||||||||||||||