Wednesday, June 29, 2011

Relative Phase Over Frequency Response

By Norman Varney

Audio enthusiasts are always concerned about frequency response. We see this data published in the specifications sheets of audio equipment, we often see it displayed graphically in reviews, and we are often interested in the frequency response of our room, etc. This is all fine, but what we should care much more about is phase.

Our experienced brain is very forgiving when it encounters missing frequencies or intensities of recognizable sounds, and it does not know the difference when frequencies are missing from unfamiliar sounds. For example, you've probably never heard an actual explosion like those in action movies, or like many, have never been in the presence of a live orchestral performance. When inexperienced, you don't know what you're missing. On the other hand, you hear the kick drum on "Billy Jean" over tiny speakers and recognize it as such. Your brain works hard to fill in the missing amplitude and low frequency information in order to make it believable. Our brain is not able to do such a great job of modeling for phase. 

What is phase?
Phase, in relationship to audio, has to do with time. Time in audio is measured in units. For example, velocity is defined in terms of length and time, or feet per second. Frequency is measured in cycles per second (abbreviated Hz.), and wavelength is measured in distance per cycle period.  A 1 kHz. tone is about 1.13 feet long and takes about 1 ms to generate, while a 100 Hz. tone is about 11.3 feet long and requires about 10 ms to produce. The standard unit of time is the second (abbreviated s). The standard clock is the Cesium-133 atom, which undergoes 9,192,631,770 oscillations per second.

You might be thinking that time and frequency are just two different mathematical ways to describe the same information in different domains, and you'd be right. However, when we start talking about more than one frequency happening simultaneously, as in a recording or playback system, we have to analyze their relationship to each other in order to determine the accuracy of what we perceive. Phase is both time and frequency dependent. Phase is the term used to describe the progress of a waveform in time relative to a starting point.


What is phase error?
Phase error results when two sound waves reach their maximum and minimum values at different times. Any degree of phase shift will cause the combined signal to be altered respectively, via the result of constructive and destructive interference.

How do we perceive relative phase distortions? 
In physics, sound is only vibration, but for the human brain, sound requires processing a lot of information in order to make sense of it as a sensation and react to it. Localization is instinctively our primary concern regarding sounds. We spatially map the location using the disparity of time (below approx.700 Hz.) and/or intensity (above approx. 700 Hz.) between our two ears. This is followed by frequency (pitch) and /or loudness, whichever wins our attention to indicate possible threat. Finally, requiring a tad more information (time), we analyze tone. We process this information a number of ways, looking for clues to discover whether the sound is friend or foe. We will look at some basic characteristics of sound as it is related to phase and human perception:

1. Amplitude. If we were to play a steady tone of say 500 Hz. on the left speaker in an anechoic (reflection-free) chamber, and then add the same tone to the right, with the same relative phase, the sound energy will have doubled in power and the result is perceived as 3 decibels louder. What if we were to delay the second tone one half cycle later in time than the first? We still have double the power, however it is 180 degrees out of phase from the first causing cancellation of the two frequencies, resulting in silence. What's happening is, as the first speaker is moving forward (compressing air molecules), the second speaker is moving backward (rarefaction of air molecules). The combination leaves the air molecules at rest.  Now you understand how phase error effects frequency response. Nature begins her sounds with a wave of compression. Electrons however, flow without regard to our human perception. You are just beginning to see how important phase is to accurate audio.



    2. Spatiality. The easiest and most drastic phase distortion that most people recognize is the confounded sound when the polarity of one stereo channel is reversed. Rather than organized in space, sounds are difficult to localize and seem disoriented, thin and hollow. Both the soundstage (the apparent physical size of the presentation) and the image (the events that take place within the soundstage) are in chaos when the two channels are 180 degrees out of phase with each other. This is an unnatural phenomenon that you feel, and your brain works hard to make sense of it for comfort, but to no avail. Less dramatic degrees of phase shift will effect spatial cues and cause sounds to be incorrectly located or wander about in apparent location and size. Spatiality cues typically occur within the first few milliseconds of the signal's introduction.

    3. Timbre. Timbre is the subjective tonality or "character" of sound. It has nothing to do with pitch or loudness per se. When hearing a flute and a violin each playing the note Middle A (440 Hz.), it is the differences in their unique attack, envelopment of harmonics (partials) and decays that distinguish them apart. This is due to not only the way an instrument is played whether: plucked, struck, blown, rubbed, etc., but also their harmonic make-up (most musical instruments posses up to twenty overtones above the fundamental), and their resonance make-up (the body of the instrument amplifies or dampens certain frequencies). Good timbre is synonymous with good fidelity, whether you are talking about a musical instrument or a hi-fi system. A cheap violin does not have the rich resonances found in a quality one, and a cheap stereo system probably won't distinguish between steel and nylon strings on a guitar, let alone the difference between Ernie Ball and D'Addario strings. It's the intensity of the overtones, during various points in time, that make these distinctions. Plomp (1970) summarized: a) Phase has maximum effect on timbre when alternate harmonics differ by 90 degrees. b) The effect of phase on timbre appears to be independent of the sound level and the spectrum. Timbre recognition occurs in about the first 20-50 ms of introduction.

     (a) The waveform of an attack transient. (b) Amplitudes of the first five harmonics of the attack transient of a 110 Hz. diapason organ pipe. (From Keeler, 1972). Notice the second harmonic develops slightly faster than the others, including the fundamental. In other woodwinds, the fundamental usually leads.

    Timbre is altered when phase is shifted. Phase distortion to the original signal confuses our brain. It is interesting to note the experiments by Berger in 1963 where he removed the first and last half seconds of 10 various band instruments and asked 30 band students to identify them. Among the confusion, only the oboe was correctly identified by more than one third of the group, eleven identified the alto saxophone as a French horn, and 25 thought the tenor saxophone was a clarinet!

    Experiment
    Though the following exercise does not follow "real world" situations, it does a great job of allowing the reader to understand and experience what happens when phase shifts alter timbre.

    While holding your hand flat with your palm facing you, say shhhhhhhhhhhhh while slowly bringing it up to your mouth. Notice how the timbre of the sound changes. You are hearing the original sound conflicting with the reflected sound off your hand. As you move your hand closer to your mouth, different frequencies (predominately around 1 kHz. - 16 kHz.) are passing through one another in opposite directions, and depending on the interval in time, or point in space you happen observe the sound, it will appear different (brighter or darker, louder or softer) at certain frequencies.


    What causes phase distortion?
    There are two types of relative phase distortions that typically occur during the recording and playback process: electrical and acoustical. And there are two causes of phase distortion: delays and repeats.

    1. Electrical. Any and all types of audio electronics will add a time delay to an applied signal, from microphones, to cables, to loudspeakers and all processors in between. Each electronic device in the signal path introduces some capacitance (stored voltage charge) and inductance (stored current charge) to the moving electrons. These inherent charges take time to develop and each signal frequency has a unique voltage and current. If the time delay is constant at all frequencies between the input and the output of the device, it is said to be phase linear or phase coherent.


    2. Acoustical. Acoustical interference occurs when room reflections cause constructive (additive) and destructive (subtractive) phase errors, as can less-than-precise speaker/listener alignment, and multi-microphone leakage. As the direct signal combines at our ears with the delayed signal(s) of itself, we experience distortion.
    a. With room reflections, and stationary listening, our brain can adapt with some "spectral compensation" to the room, especially in the higher frequencies. However, reflections that are within -15dB of the direct sound will definitely cause audible phase anomalies.
    b. Ideally, each speaker voice coil should be the same distance to the listening position so that the signals from each arrive together. When they are not aligned, the relative signal arrival times are different, causing change to the sound, and to the polar response (directivity) of the speaker. Note that the more off-axis a listener is, the more time incoherency is increased. Note also that good designs take into account cross-over network phase and delays, and that even a 5 us change can be audible.
    c. Two microphones, each in a different location, but both picking up similar information can cause tonality errors. For example, a snare top head and bottom head mic both picking up the high-hat, or the bottom head mic picking up the direct sound with reflected sound from the floor.



    What can you do to reduce phase errors?
    There are several things one can do, even if you don't have sophisticated test equipment or knowledge:
    1. Train your ears. Listen to unamplified music for reference. Enjoy the richness of harmonic content, the spatial imaging, the attack, envelopment and decay of individual sounds.
    2. Confirm that all amplifier/speaker channels are the same polarity.
    3. Confirm that all speaker drivers are the correct polarity. Placing a 9 volt battery across the speaker cable leads should push all drivers forward in nearly all speaker designs.
    4. Investigate interconnects, speaker cables and speakers that boast about energy efficiency, time/phase alignment, etc.
    5. Do your best to locate the speaker/listening position for smoothest room mode response in the room.
    6. Confirm that each speaker is the same distance to the listening position.
    7. Treat first order reflection points in the room with absorption or diffusion. This can be done with the "mirror trick". Treat the locations with at least a 2' area to cover frequencies down to about 500 Hz.

    Conclusion
    As with many blogs, a book could be written about the subject. Phase is easily one of them. Only scratching the surface, there are many sub-topics of phase effects that I did not include, such as pitch, resonance, ringing, beat frequencies, etc.

    Time delay spectrometry has only been around since the late 1960's. Prior to that, we didn't have the computer processing required to analyze the relationships between time, energy and frequency. This may be why we are so concerned and familiar with frequency response. Phase distortion is the primary reason why one piece of equipment sounds different from another. It is also the primary reason why most people are denied the full potential of their audio investment and cannot enjoy the full experience created by the artist.  
     

    In regards to relative phase perceptibility over frequency response, most any piece of good audio equipment today will offer good frequency response, but most do not have good phase response. The problem for the end-user is integrating synergistic system components, setting them up properly both physically and electronically, and controlling room acoustics. Noticing minor phase errors in a typically reverberant room will be difficult because of the lack of resolution available from the room. Controlling the acoustics properly is like discovering the deep sea. You probably have no idea of all the cool stuff that is below the surface.


    Thursday, March 3, 2011

    5 Reasons Off-center Room Positioning is a Bad Idea

    By Norman Varney
       Symmetry of the audio scene, especially the front horizontal plane, is important to accurate reproduction. We want to place ourselves in the middle of the left and right image in order to hear the proper soundstage. If we don't balance levels correctly, spatial cues, frequency response, low-level details, etc. become skewed because the energy on one side of us is louder than the other. This is true with headphones, but is even more problematic when sound is introduced to a room. Lately, I've been seeing a lot of designs that are incorporating off-center speaker/listener arrangements in the room. The idea is to avoid the room's fundamental width cancellation node by moving away from it. This is not practical for high fidelity.


      Axial room modes in a rectangular room are fairly predictable using simple math. Since room modes are dictated by room dimensions, we can calculate what frequencies will live where in the room. We want to avoid coupling the woofers and listeners with the existing first, second and third-order modes whenever possible, as they are the most energetic. Avoid placing woofers in antinodes (pressure peaks), and listeners in both nodes and antinodes. Placing speakers in an antinode will excite it, resulting in that frequency (and its harmonics) sounding louder than they should.  Placing a speaker in a node (null) will attenuate that mode, which at times can be useful. Placing a listener in an antinode results in the mode sounding too loud. Placing a listener in a node results in the mode sounding too soft. There is always an optimum position for the speakers/listeners in a room to deliver the best soundstage and bass response. 


      Symmetry is important. Placing the speakers and listeners off-center in the room to avoid the fundamental width mode is not a good idea. Here are some reasons why off-center positioning is not good practice:

    1. The fundamental (f1) mode wavelengths are too large to move away from. The fundamental wave supported in a 15' wide room is about 30' long (38 Hz.). The longer the dimensions, the longer the lowest wavelength. You would have to move off-center about 3.75' to smooth out a 38 Hz. null.
    2. By doing so, you’ll just end up in another mode. In a 15' wide room, 3.75' off-center, you'll find f3 (113Hz.) at its peak.
    3. By doing so, you’ll end up too close to the side wall, which will cause timing differences between your left and right ears, resulting in severe spatial skewing. 
    4. By doing so, you’ll end up too close to the side wall, which will cause energy differences between your left and right ears, resulting in resolution loss and inequality.
    5. By doing so, you’ll end up too close to the side wall, which will cause frequency differences between your left and right ears, resulting in severe timbre skewing and inequality.

      Let's look at what happens at these low frequency room modes. If we took an instantaneous time snapshot (1/75th of a second) of the first-order (f1) width mode in a room 15’ wide (38 Hz.), we would see a positive pressure point to our left, and null in the middle of the room, and a negative pressure point to our right. At the same instant, the second-order (f2) mode (75 Hz.), which is half the length of the first, would show a positive peak at the left wall followed by a null (located about 3.75’ from the left wall), a positive peak in the middle of the room, and a null (located 3.75’ from the right wall), followed by a positive peak at the right wall. We want to avoid the third-order (f3) mode as well (113 Hz.). You would have to move 3-4’ to one side before you would notice any appreciable frequency smoothing of the first-order mode, which moves us into to the f3 antinode at 113 Hz. This particular frequency is contained in nearly all music and dialog recordings. Not a good move (see Fig. 1).

    In summary, we must place the audio footprint center of the side walls and settle for the rare, problematic bass note, over distorting all frequencies, all the time. Placing the auditory scene symmetrically between the left and right walls provides optimum dynamics, tonality, imaging, spacial cues and low-level resolution.


    Enhanced by Zemanta

    Thursday, February 17, 2011

    10 Reasons Why Frequency Averaging is NOT a Good Idea

    By Norman Varney

    The idea of averaging multi-channel sound for all listener positions is not a good one. There are many audio myths out there. This one has many professionals fooled. It is a popular practice that should be better understood before it is applied, especially to small, critical listening spaces such as home cinemas, production suites, etc.

     

    I am not dissing the use of equalizers. They are a powerful tool and can be used for many reasons, but in this instance, I am cautioning their use to average the frequency response to be received over the entire listening zone. This practice is almost a given in the home theater arena and should be reexamined. There are more reasons why frequency response spatial averaging is not a good solution. I am only going to cover the main reasons:

     

    1) It won't average the time domain at more than one location.

    2) It won't average energy levels at more than one location.

    3) It won't average the reverberation times at more than one location.

    4) Each seat will still have a unique sound even with frequency averaging.

    5) Only one seat can be calibrated for audio nirvana at a time. This is because multi-channel sound can only converge at the same time, energy and frequency response, at a single point in space (See Fig. 5).  This location should be the primary seat, which should be located in the middle of the listening zone, so that more listeners are closer to the target.

    6) If we average the frequency response for all seats, all seats suffer (See fig. 6).

    7) Equalizers can only change the the frequency response of the signal being fed to the speaker(s). They do not respond to room interaction. Once the signal leaves the speaker, the room takes control of the sound energies.

    8) Equalizers are good for reducing peaks and poor at increasing dips.

    9) The differences between seats is not small. They can easily differ by 30 dB SPL.  As the number of seats increase, so does the variance, as do rooms with seats close to boundaries and/or rooms with particularly bad modes.

    10) Selecting a "room curve" for spatial averaging, guarantees the original signal is lost.

     

    The point in the room where time, energy and frequency converges should be the center of the listening space. Any other seat will be compromised to a degree regardless. Why compromise every seat? As depicted in Fig. 5 & 6, more seats will exhibit better sound without averaging frequency response. In addition, audio nirvana can be experienced in the primary seat.

     

    The idea that "global equalization" equates to "great sound at every seat", or even that the sweet spot becomes bigger is not accurate. Applying this scheme means degrading the sound for everyone and lessening the experience otherwise possible for the money seat.  

     

    Sometimes poor room mode distribution may call for some electronic equalization. A 1/20th octave band parametric equalizer could be introduced to smooth these low frequencies (around 300 Hz. and below). Remember, we are changing the frequency response of the signal driving the speaker, and are trying to compensate for how the room is reacting to it from our personal point of view. 

     

    Room mode anomalies can mean a difference of 30 dB between a peak and valley. An electronic equalizer can do a good job of reducing peaks, but can not offer more than about 6 dB of gain for a coincident dip. 

     

    Note that most of the equalizers used for this process use steady-state measurements with a microphone collecting sounds from every direction in the room. It sums them together for analysis and then a programed solution is applied. This is not the way humans process sound with two ears and a brain. Even fancy digital signal processors (DSP) with time domain correction cannot compensate for room reflections, etc. because they cannot separate where the sound is coming from. As a side note, they also sacrifice frequency resolution in order to analyze time.

     

    In summary, electronic room equalization, ideally, is used when passive corrective means are not possible, or in conjunction with passive means. Ideally, this equalization should be of the high resolution, high Q, parametric type, and only be used to address room mode problems around 300 Hz. and below. This is assuming that the speakers used are accurate to begin with. If they are, global equalization across the audible spectrum will likely do more harm than good for all seats. 

     

    I've listed ten reasons why spatial frequency averaging is not a good idea. There are many more. Can you think of some of them? Please comment in the box below.



     

    Thursday, February 10, 2011

    How Bad Can a 1% Air Gap Be to Noise Control?


    by Harry Alter and Norman Varney 

    Noise control is a two way street. You may have spent considerable expense on the design and materials of a wall, ceiling or floor system to keep noise from escaping or entering the space. However, you may not realize the impact a tiny hole can have on the entire partition's performance. Most people know that the door is the weak spot in a wall system. You can have a high sound transmission class (STC) rated wall system cut to half the rating if you don't incorporate an acoustical door, or not have one properly installed.

    There are two primary means that sound energy can travel through walls, floors and ceilings; vibrations traveling through solid materials such as gypsum, sheathing, studs or joists are called structure-borne vibrations, and vibrations traveling through air, framing cavities and unsealed penetrations, seams or gaps are called air-borne vibrations. Both vibration paths play an important role in determining how well the partition assembly will reduce the transmission of sound through it, and a "systems" approach must be in its design to appropriately address the associated sound energies. A systems approach would incorporate a combination of blocking, breaking, absorbing and/or isolating the energy at the source, along its path(s) and/or at the receiver. We are covering just one of these aspects of noise control in this particular blog.

    Air filtration and sound penetration through walls, ceilings and floors occur as one in the same. If air can penetrate a partition, then so can sound. In fact, it takes very little air leakage to cause significant sound leakage. For example, an opening or crack 1/100th of 1% of a total wall's surface area can reduce the sound transmission loss (TL) of a wall from 50 to 39 dB. That's an 11 dB drop in noise control performance. Likewise, a partition designed to achieve a TL of 40 dB would be reduced to approximately 30 dB (a 10 dB drop) with only 1/10th of 1% air leakage area to wall area. Note that the 10 dB drop in the poorer assembly would be perceived by the average person as twice as loud as the better assembly.

    The above graph illustrates how openings and cracks can affect the TL (and subsequent STC) of a partition assembly. The horizontal axis indicates the design or desired performance of the assembly. The vertical axis indicates the resultant TL based on the % of surface area air/sound leakage. From the graph, the level of noise control performance will not increase beyond a certain level based on the size of the unsealed air gap. As a result, sealing air gaps reduces sound (noise) transmission through partitions. Less air penetration equals less sound penetration. The beauty is in the details.




    Saturday, January 29, 2011

    I'm happy to hear that I should still enjoy music if I contract Alzheimer's. Check out author Daniel Levitin at http://ow.ly/3MD6A.

    Friday, January 28, 2011

    Tip for those introducing computer-music files to their playback systems- at minimum, plug it into a different circuit & typically left off.

    Tuesday, January 18, 2011

    Constructing a suspension mount today to perform a series of low-level, broad-band vibration tests.