Loss of Stereo Image

Author: Jürgen Herre




Background information: Spatial perception
Although spatial perception is far from being fully understood, it is known that directional localization of sounds depends on the evaluation of so-called spatial cues by the human auditory system (see e.g. Blauert's book on spatial hearing [1]), the most important cues being

  • interaural level differences (i.e. the differences in levels received by both ears)
  • interaural phase differences (i.e. the differences in signal phase received by both ears)
  • preservation of fine temporal signal envelope structure

Consequently, the fidelity of the stereo image of a coded signal depends on the coder's ability to preserve these critical cues appropriately.

Background information: Joint Stereo Coding
For coding of high quality stereophonic (or multi-channel) audio signals at low bit rates, joint coding techniques have proven to be extremely valuable. On one hand they provide mechanisms to account for binaural psychoacoustic effects, on the other hand the required bit rate for the stereophonic signals may be reduced significantly below the bit rate for separate coding of the input channels.
Currently, the most common joint stereo coding techniques are Mid/Side (M/S) stereo coding [2] and Intensity Stereo coding [3] [4]. While the first method can account for binaural masking effects and achieve a certain amount of signal-dependent gain, the intensity stereo method provides a high potential for bit saving. Coders based on the intensity stereo principle have been described in the past for stereophonic and multi-channel coding under various names (e.g. "dynamic crosstalk" [5], "channel coupling" [6]).
Intensity stereo exploits the fact that the perception of high frequency sound components (e.g. above 4 kHz) mainly relies on the analysis of their energy-time envelopes [1] rather than the waveform itself. Thus, it is assumed sufficient to code the envelope of such a signal instead of its waveform. This is done by transmitting one common set of spectral coefficients ("carrier signal") that is shared among several audio channels instead of separate sets for each particular one. In the decoder, the carrier signal is scaled independently for each signal channel to match its original average envelope (or signal energy) for the respective coder frame. The scaling information is calculated and transmitted once for each group of spectral coefficients (scalefactor band). Effectively, the stereo image is recreated at the decoder side by a pan-pot-like operation for each spectral coder band.



Figure 1: Basic principle of intensity stereo coding

Some typical stereo image loss problems:
As a consequence of the intensity stereo coding / decoding process, all output signals reconstructed from a single carrier are scaled versions of each other, i.e. they have the same envelope fine structure for the duration of the coded block (e.g. 10-20 ms). This does not present a major problem for stationary signals or signals having similar envelope fine structures in the intensity stereo coded channels.
For transient signals with dissimilar envelopes in different channels, however, the original distribution of the envelope onsets between the coded channels cannot be recovered. Figures 2 and 3 show an example for this constellation: In a stereophonic recording of an applauding audience, the individual envelopes will be very different in the right and left channel due to the distinct clapping events happening at different times in both channels (see Figure 2 for left and right channel envelopes).

HF Envelopes

Figure 2: High frequency envelope structures of left and right original channel signals. Excerpt from "applause" item


After the intensity stereo encoding / decoding process, the fine time structure of the signals is mostly the same in both channels as can be seen in Figure 3. In particular, there is a structural "cross-talk" between the channels, such that perceptually important signal onsets propagate to the other opposite channel (e.g. L->R at 15 ms, R->L at 47 ms, L->R at 57 ms, L->R at 70 ms).




Figure 2: High frequency envelope structures of left and right channel signals after Intensity Stereo encoding/decoding (red). Envelopes of the original signals are shown in dotted yellow lines. Excerpt from "applause" item

Consequently, the stereo image quality of the intensity stereo coded / decoded signal will decrease significantly in such cases. The spatial impression tends to narrow down and the perceived stereo image collapses into the center position. For critical signals, like the applauding audience example, the achieved quality cannot be considered as acceptable anymore.

Sound examples
The following sound excerpts illustrate the discussed effects:

Applause Original stereo sound excerpt:
A wide stereo image
Applause 6k Intensity stereo encoded/decoded
(starting from 6 kHz)
Applause 4k Intensity stereo encoded/decoded 
(starting from 4 kHz)
Applause 2k Intensity stereo encoded/decoded 
(starting from 2 kHz)
Applause_1k Intensity stereo encoded/decoded 
(starting from 1 kHz)

The provided example sound files demonstrate the original applause recording as well as three sound examples with increasing deficiencies in stereo imaging quality, as would be produced by a coder without a proper control of the intensity stereo coding mechanism. Please observe the increasing loss of people applauding in the outer left and outer right seats as well as the overall lack of spatial impression and distinct reproduction of the single clap events.

References:
[1] J. Blauert, "Spatial Hearing", MIT Press, 1983

[2] J. D. Johnston, A. J. Ferreira: "Sum-Difference Stereo Transform Coding", IEEE ICASSP 1992, pp. 569-571

[3] R.G.v.d. Waal, R.N.J. Veldhuis, "Subband Coding of Stereophonic Digital Audio Signals", IEEE ICASSP 1991, pp. 3601 - 3604

[4] J. Herre, K. Brandenburg, D. Lederer, "Intensity Stereo Coding", 96th AES Convention, Amsterdam 1994, Preprint #3799

[5] G. Stoll, G. Theile, S. Nielsen, A. Silzle, M. Link, R. Sedlmayer, A. Breford, "Extension of ISO/MPEG-Audio Layer II to
Multi-Channel Coding: The Future Standard for Broadcasting, Telecommunication, and Multimedia Applications", presented at the 94th AES Convention, Berlin 1994, Preprint # 3550

[5] Mark Davis, "The AC-3 Multichannel Coder", 95th AES Convention, New York October 1993, Preprint # 3774