|
and their variation in time ("Birdies") Author: Antonio Pena Co-Author: Enrique Alexandre Background: pattern detection and timbre A perceptual audio codec includes, by definition, a module implementing performing a model of the Human Audition System [4]. Most basic codecs derive a masking threshold estimation to determine the highest level of noise that will be imperceptible allowed at a certain time-frequency (t-f) positions., in order to be not perceptible. Other measures such as loudness or pitch modelling may also be part of the model. But These kinds of models basically resemble the Auditory Interface (external, middle and inner ear), without considering the higher level processing to be carried out by the brain. These high- level processes seem to divide an entire sound input into a collection of independent items (auditory objects), and organize these items into several sets (streams) [2]: For example, a tune containing both a violin and a flute playing at the same time would contain many audio objects (every single note played) and two a couple of streams, one for the notes played by the violin and another for the notes coming from the flute.; so, Every object would be assigned to a certain given stream (auditory streaming) by the cognitive processing. Human perception (audition, vision, etc....) seems to be driven by some basic rules pointed out in some of the Gestalt Psychology texts (Germany, beginning of the twentieth century) [3] . These are the most relevant for the next examples:- Similarity: elements are grouped if similar. - Good continuation: a smooth change keeps the elements grouped in the same set. - Common fate: if some elements change at the same time, they are grouped in the same set. - Continuity: even if some parts of an acoustic event are lost, previous memory may fulfil any existing gap. - Selective attention: perception can be tuned to a particular stream (analytic) or relaxed in order to perceive something as a single compact set (synthetic). This behaviour probably makes use of both physical (tracking of frequency components, time onsets, temporal and frequency envelopes, etc....) and psychological (presence of other sources, previous knowledge, images or other events associated to the acoustic excerpt, etc....) cues to process all this information together. Understanding timbre as the whole set of characteristics that remain once pitch, loudness and duration are extracted, then the auditory streaming assembles some of the t-f components into a single auditory object, taking into account some Time-Frequency-relationships between them. It is especially remarkable is that isolated components create new timbres and, so, new perceived objects; on the other hand, as commented above, continuity preserves timbre and keeps tracking of the temporal evolution of components. Background: bit assignment and coding bands Sound examples: timbre vs. some typical bandwidth limitations
* Funky sound: please note the cymbals set when listening to this example. As you will notice, the bandwidth truncation rapidly blurs the characteristic timbre and the cymbals almost disappear as a perceived instrument when the filtering is more severe. On the other hand, bass and keyboards do not suffer so much with the reduction of frequency components.
It is not difficult to order these funky sounds in a decreasing bandwidth
ordering but, what about the previous ones? Try to order these Glockenspiel
sounds: GlockenA,
GlockenB, GlockenC,
GlockenD. Sound examples: some typical bandwidth variations through time (the
birdies artifact) Now consider a natural audio signal that has been coded without any bandwidth
control and sosuch that strange changes in the presence of some bands
may appear. There are three different coded versions, from better to worse,
where you clearly perceive some strange high- frequency effects that appear
and dissappear, with no specific connection withto the original sound.
Some of these artifacts seem to be in the upper portion of the spectrum
and in somethese cases a bandwidth clipping is recommended to avoid theis
presencedistortion (Figure 2.a,
Figure 2.band
Figure 2.c)., In other cases, however, the artifacts but sometimes
are within the synthetic spectrum and bandwidth limitation is not a solution
(Figure 3.aand
Figure 3.b). Listen to the previous files low pass filtered down to
8 KkHz and observe that not every coded file has become a birdies-free
sound.
References [1] J.Beerends and J.Stemerdink, "Modelling a cognitive aspect in
the measurement of the quality of music codecs", in Preprint 3800,
96th AES Conv., Amsterdam, 1994.
|
|||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||