Artifacts in frequency
and their variation in time ("Birdies")


Author: Antonio Pena
Co-Author: Enrique Alexandre



Background: pattern detection and timbre

A perceptual audio codec includes, by definition, a module implementing performing a model of the Human Audition System [4]. Most basic codecs derive a masking threshold estimation to determine the highest level of noise that will be imperceptible allowed at a certain time-frequency (t-f) positions., in order to be not perceptible. Other measures such as loudness or pitch modelling may also be part of the model. But These kinds of models basically resemble the Auditory Interface (external, middle and inner ear), without considering the higher level processing to be carried out by the brain.

These high- level processes seem to divide an entire sound input into a collection of independent items (auditory objects), and organize these items into several sets (streams) [2]:  For example, a tune containing both a violin and a flute playing at the same time would contain many audio objects (every single note played) and two a couple of streams, one for the notes played by the violin and another for the notes coming from the flute.; so, Every object would be assigned to a certain given stream (auditory streaming) by the cognitive processing.

Human perception (audition, vision, etc....) seems to be driven by some basic rules pointed out in some of the Gestalt Psychology texts (Germany, beginning of the twentieth century) [3] . These are the most relevant for the next examples:

- Similarity: elements are grouped if similar.
- Good continuation: a smooth change keeps the elements grouped in the same set.
- Common fate: if some elements change at the same time, they are grouped in the same set.
- Continuity: even if some parts of an acoustic event are lost, previous memory may fulfil any existing gap.
- Selective attention: perception can be tuned to a particular stream (analytic) or relaxed in order to perceive something as a single compact set (synthetic).

This behaviour probably makes use of both physical (tracking of frequency components, time onsets, temporal and frequency envelopes, etc....) and psychological (presence of other sources, previous knowledge, images or other events associated to the acoustic excerpt, etc....) cues to process all this information together.

Understanding timbre as the whole set of characteristics that remain once pitch, loudness and duration are extracted, then the auditory streaming assembles some of the t-f components into a single auditory object, taking into account some Time-Frequency-relationships between them. It is especially remarkable is that isolated components create new timbres and, so, new perceived objects; on the other hand, as commented above, continuity preserves timbre and keeps tracking of the temporal evolution of components.

Background: bit assignment and coding bands
The basic core mechanism of an audio coder [5] is a set of different quantizers that can be chosen to process a given area of frequency (subband, coding band). The selection comes from a bit assignment routine that decides which is the quantizer to use or, in other words, which is the amount of quantization noise to be injected into the synthetic signal. This assignment is usually time-varying (frame-by-frame) and uses perceptual information. One of these available quantizers is the null one (zero bits assigned and, so, the quantized components are lost).
 

Sound examples: timbre vs. some typical bandwidth limitations
Timbre characteristics depend strongly on frequency, in addition toas well as time structure, but not every sound source has the same dependence. Some sound sources hold their main characteristics in the first 4 or 5 kHz (e.g. speech), and most of their sounds are clearly recognizable with such a limited bandwidth (e.g. telephone). Of course high quality speech needs a higher bandwidth, such as 7 or 8 kHz, but there are other types of sounds that need a much higher bandwidth, becoming unrecognizable, or just disappearing, if a truncated bandwidth is used. Let us listen to the first two examples:

* Glockenspiel: you may find this set of notes is lowpass filtered through in the next excerpts. It is a percussive sound but its tonal nature probably preserves part of the timbre components in spite of bandwidth truncation (listen to the last four notes even at 4 kHz low pass filtering).
 
Glocken Original (bw: 20 kHz )
Glocken_12k 12 kHz low pass filtered
Glocken_8k    8 kHz low pass filtered
Glocken_4k    4 kHz low pass filtered

* Funky sound: please note the cymbals set when listening to this example. As you will notice, the bandwidth truncation rapidly blurs the characteristic timbre and the cymbals almost disappear as a perceived instrument when the filtering is more severe. On the other hand, bass and keyboards do not suffer so much with the reduction of frequency components.
 
Funky Original (bw: 20 kHz )
Funky_12k 12 kHz low pass filtered
Funky_8k    8 kHz low pass filtered
Funky_4k    4 kHz low pass filtered

It is not difficult to order these funky sounds in a decreasing bandwidth ordering but, what about the previous ones? Try to order these Glockenspiel sounds:  GlockenA, GlockenB, GlockenC, GlockenD.

Sound examples: some typical bandwidth variations through time (the birdies artifact)
Consider now the synthetic signal Chord 440(Figure 1.a) as the next example. It contains a fundamental "A" note (440 Hz) plus four harmonics. The five tones fuse together in a perfect octave consonance, that is, a single perceived stream. When suppressing the 1320 Hz harmonic, the four remaining tones fuse together in a perfect octave consonance, but with a different timbre: Chord 440 without 1320 Hz tone(Figure 1.b). Moreover, as the energy from the cancelled harmonic is not present, perceived loudness has reduced but just one stream is still clearly perceived. But let us In the next example, we variably suppress the 1320 Hz harmonic here and there, making it appearing and disappearing from the sound: Chord 440 with time-varying 1320 Hz time-varying tone(Figure 1.c) . Now the 1320 Hz tone outstands out as a distinct item, clearly separated from the remaining four stationary tones that form the chord. When listening to the last 3 seconds, you will probably perceive how the tone fuses again within the consonant chord once it becomes again stationary again.

The most used perceptual measure in audio coding is the masking threshold, which is usually computed frame by frame. The fact is that slight variations of the masked threshold from one frame to the next, however, may lead to very different bit assignments when the bit rates is low., and as a result, some groups of spectral coefficients may appear and disappear. This spurious energy constitutes several auditory objects, which are different from the main one and are thus clearly perceived. These kind of artifacts, known as birdies, have been reported both for the tuning of audio codecs [6] and for objective perceptual assessment methods [1].

Now consider a natural audio signal that has been coded without any bandwidth control and sosuch that strange changes in the presence of some bands may appear. There are three different coded versions, from better to worse, where you clearly perceive some strange high- frequency effects that appear and dissappear, with no specific connection withto the original sound.

Gilmour Original (bw: 20 kHz )
Gilmour1 Slightly distorted decoded signal
Gilmour2 Clearly distorted decoded signal
Gilmour3 Severely distorted decoded signal

Some of these artifacts seem to be in the upper portion of the spectrum and in somethese cases a bandwidth clipping is recommended to avoid theis presencedistortion (Figure 2.a, Figure 2.band Figure 2.c)., In other cases, however, the artifacts but sometimes are within the synthetic spectrum and bandwidth limitation is not a solution (Figure 3.aand Figure 3.b). Listen to the previous files low pass filtered down to 8 KkHz and observe that not every coded file has become a birdies-free sound.

Gilmour_8k 8 kHz low pass filtered original
Gilmour1_8k Almost birdies-free
Gilmour2_8k Slightly worse decoded signal
Gilmour3_8k Severely distorted decoded signal

  References

[1] J.Beerends and J.Stemerdink, "Modelling a cognitive aspect in the measurement of the quality of music codecs", in Preprint 3800, 96th AES Conv., Amsterdam, 1994.
[2] Albert S. Bregman, "Auditory Scene Analysis: the Perceptual Organization of Sound", MIT Press, 1990.
[3] K. Koffka, "Principles of Gestalt Psychology", Kegan Paul, Trench, Trubner & Co., Ltd., London, 1936.
[4]Brian C.J. Moore, "An Introduction to the Psychology of Hearing", Academic Press, 1989.
[5]T. Painter and A. Spanias, "Perceptual coding of digital audio", Proc. of the IEEE, vol. 88, no.4, pp. 451-513, April 2000.
[6] A.S. Pena, "Técnicas de modelado psicoacústico aplicadas a la codificación de audio de muy alta calidad", PhD thesis (in Spanish), Universidad Politécnica de Madrid, 1994.