Essay

The Emperor's New Binaural

Dolby calls its headphone render binaural. The spectrum, the interaural timing, and the renderer binary all describe an amplitude panner with a reverb plugin.

The audio industry runs on a specific kind of lie: the kind that lets engineers build products for people who cannot differentiate a sample rate from a coffee roast. The lie needs a name that sounds scientific but stays vague enough to mean whatever the marketing department requires it to mean at any given moment. Dolby has perfected this art.

Their latest term of art is binaural. When you see it in the Dolby Atmos Renderer, in the ADM metadata, or in some press release written by people who have never coded a single sample, it implies the pinnacle of spatial audio science. The suggestion is that Dolby’s algorithm models how sound interacts with your head, your pinnae, the shape of your outer ear, and the folds of your nasal concha to deliver a true, scientifically accurate 3D audio experience. It’s a beautiful story. It’s also almost entirely false.

The evidence against the claim converges from two directions: what comes out of the system, and the binary code itself. Both tell the same story. Dolby binaural rendering is amplitude panning plus room reverb with a thin magnitude-domain HRTF component. Binaural is the word on the box. What ships in the box is much simpler.

~1 dBSpectral match between Dolby's “binaural” output and a plain amplitude panner.
1.6%Frames carrying >100 µs of interaural delay. Panner: 0.6%. Full HRTF: 9.1%.
0Impulse responses in 4.6 MB of coefficient tables inside the renderer binary.
Methods, full tables, and limitations → read the paper (PDF)

The Spectral Evidence

Start with the obvious measurement: treat the Dolby Atmos Renderer as a black box and compare its output against a known reference. To do this properly, I needed a ground truth binaural renderer that actually does full per-object HRTF convolution. I built this twice, once in JavaScript and once in Rust, using Steam Audio HRTF sets and various SOFA files to ensure the language and library implementation didn’t introduce artifacts. The input was ADM-BWF production masters: real Atmos content with objects positioned in 3D space.

The comparison was stark: Dolby’s official binaural re-render versus my full HRTF convolution. The results were not close. Full HRTF convolution produces a consistent 10+ dB spectral dip compared to Dolby’s output, with pinna-specific notches at ~6.5 kHz (10+ dB) and 9-10 kHz (up to 15 dB). These notches are what give your brain the spectral cues necessary to actually locate sound in 3D space. They persist across every HRTF dataset, both implementations, and every configuration.

Log-frequency spectrum from 20 Hz to 20 kHz. Dolby's re-render (gold) stays smooth above 2 kHz; the full-HRTF render (red) falls 10 to 20 dB below it, with deep notches in the shaded 5 to 8 kHz region and again near 9 to 10 kHz.
Fig. 1Dolby re-render vs. full per-object HRTF, same ADM-BWF master. This is what real pinna filtering does to a dense Atmos mix. Dolby's output shows none of it.
Zoom on 4 to 10 kHz. Dolby (gold) and the 15% blend (green) run together near minus 65 to minus 70 dB. Full HRTF (red) sits 10 to 20 dB lower, with a deep notch below 6 kHz and another near 9.5 kHz.
Fig. 2The pinna band, 4–10 kHz. Full HRTF carves the notches your brain uses for elevation. Dolby and a 15% blend sail straight over them.

The second approach was to see what happens if I turn off the spatial processing entirely and just do amplitude panning. The result was nearly identical to Dolby’s binaural output. The spectral match was within ~1 dB. Dolby’s output shows only a small residual dip at 6.5 kHz that pure panning lacks. A ~1 dB match to a dumb panner is a confession.

Dolby's re-render (gold) and a pure amplitude-panning render (cyan) overlaid from 20 Hz to 20 kHz. The two traces lie on top of each other across the whole band.
Fig. 3Switch the HRTF off completely. Pure amplitude panning lands on top of Dolby's “binaural” render from 20 Hz to 20 kHz.

A ~1 dB match to a dumb panner is a confession.

To find the blend, I tested a 15% HRTF / 85% panning mix on the reference track and matched Dolby’s output almost exactly. Across five different tracks, the ratio refused to hold still. A fixed HRTF convolution has one ratio. A heuristic tracks whatever the panner is doing. Dolby’s blend drifts because the HRTF component rides on the panning engine instead of driving it.

Four renders of the reference track overlaid: Dolby (gold), full HRTF (red), panning only (cyan), and a 15% HRTF blend (green). Gold, cyan and green form one band; red is the only trace that departs, falling far below above 2 kHz.
Fig. 4All four renders of the reference track. The 15% blend hides the panning-only trace almost completely, and both sit on Dolby. Only full HRTF breaks away.
Four spectrum panels, one per Atmos item: Greece, Accidental Effects, Fluid, Dialogo Interno. In each, Dolby (gold) and our renderer (green) overlap. RMS differences are 0.42, 0.52, 0.45 and 0.53 dB.
Fig. 5The same comparison on four more Atmos items of different genre and density. Dolby vs. our renderer: RMS difference 0.42–0.53 dB on every track. Dolby's timbre tracks the panner.

The Temporal Lie

Spectral analysis tells you what frequencies are present. It doesn’t tell you how they are arranged in time, which is where the real lie begins. The key spatial cue for binaural hearing is Interaural Time Difference (ITD): the tiny delay between sound hitting one ear versus the other. A sound at 90 degrees left should reach the left ear ~660 microseconds before the right. This delay is the most basic mechanism by which humans locate sound.

To test this, I built a controlled ceiling test. I reconstructed an ADM-BWF: 7.1.2 layout, BS.2076 axml chunk, and a broadband noise probe at known azimuths. I then ran this through my renderer at full HRTF and at pure panning. I measured the ITD using a hardened phase-domain method: coherence-gated normalized cross-correlation over 85 ms frames in the 2-12 kHz band. This approach was chosen because earlier GCC-PHAT energy-gated metrics produced spurious lags on reverberant frames, which would have clouded the results.

The ceiling results were physiologically correct. A probe at 90 degrees left produced an ITD of -667 microseconds and an ILD of +15.7 dB. At 45 degrees left, the ITD was -417 microseconds. These are real, physical cues.

AzimuthRenderITDILD
90° LFull HRTF−667 µs+15.7 dB
90° LPanning0 µs+32.9 dB
90° RFull HRTF+604 µs−17.2 dB
90° RPanning0 µs−32.9 dB
45° LFull HRTF−417 µs+11.7 dB
45° LPanning0 µs+12.6 dB
Bar chart of interaural time difference for a broadband probe. Full HRTF: minus 667 µs at 90 degrees left, minus 417 µs at 45 degrees left, plus 604 µs at 90 degrees right. Panning: 0 µs at every azimuth.
Fig. 6The controlled ceiling. Full HRTF reaches the ~660 µs head-geometry maximum. Panning produces exactly zero, by construction. The reference renderer is not the weak link.
Panning · 90° left
Full HRTF · 90° left
ListenHeadphones on. The same broadband probe at 90° left, six seconds each, straight out of the reference renderer. Panning moves the sound by level alone. Full HRTF adds the interaural delay and pinna filtering from the table above.

Across four validation tracks, I compared the fraction of frames where the ITD exceeded 100 microseconds. For pure panning, this occurred in 0.6% of frames. For full HRTF convolution, it occurred in 9.1% of frames. This is expected: panning produces no interaural delay by construction, while delay is intrinsic to HRTF convolution. Dolby’s results? 1.6% of frames. Dolby’s ITD is tracking the panning floor. It is effectively zero for any coherent signal.

Grouped bar chart per track (Greece, Accidental, Fluid, Dialogo) of frames with ITD above 100 µs. Full HRTF reaches 8 to 14 percent on three tracks; Dolby stays below 1 percent on three tracks and about 5 percent on Dialogo; panning is near zero except Dialogo at about 2.4 percent. Means: panning 0.6, Dolby 1.6, full HRTF 9.1 percent.
Fig. 7Frames with |ITD| > 100 µs, per track. Dolby hugs the panning floor. Full HRTF is in a different league.

Dolby’s ITD position between the panning floor and the full HRTF ceiling sits at ~19%, the signature of a system that does not model timing cues. You can see it in the decorrelation data. A pure EQ on a panned signal cannot reduce interaural coherence: you apply the same filter to both channels and you get scaled copies. The interaural coherence remains at 1.0. Dolby’s output, however, shows a measurable drop in coherence. This proves that Dolby’s processing is doing something more than simple panning. The ITD results identify the mechanism by elimination: magnitude-domain processing that happens to decorrelate the channels.

Frequency-resolved coherence analysis shows that Dolby’s added decorrelation is split ~47% below 2 kHz and ~53% in the 5-12 kHz band. The presence of significant decorrelation at 200-500 Hz is the smoking gun. At these frequencies, the human pinna is physically inert. It cannot possibly be the source of decorrelation. The only plausible explanation is that Dolby is adding broadband reverb to simulate distance and space, which incidentally reduces interaural coherence.

Interaural coherence versus frequency, mean of four tracks. Panning (blue) stays near 0.7 to 0.85. Dolby (gold) runs consistently below panning across the whole band, including the shaded region below 500 Hz where the pinna cannot act. Full HRTF (red) swings widely.
Fig. 8Interaural coherence vs. frequency. Dolby drops below panning everywhere, including the shaded band under 500 Hz, where no pinna can touch the signal. That is reverb, not a head-related transfer function.

The synthesis is clear: Dolby matches panning in long-term magnitude and in interaural timing, then decorrelates the ears via reverb. It’s an anisotropic processing model. Panning in time, reverberation in interaural difference, and almost nothing in pinna-grade HRTF.

Horizontal bars placing Dolby between panning at 0 percent and full HRTF at 100 percent on three cues: long-term magnitude 5 percent, interaural delay 19 percent, interaural decorrelation 72 percent.
Fig. 9No single dial describes it. On a 0% = panning → 100% = full HRTF scale, Dolby sits at 5% on timbre, 19% on timing, and 72% on decorrelation, which Fig. 8 traces to reverb.

Panning in time, reverberation in interaural difference, and almost nothing in pinna-grade HRTF.

The Binary Confession

When the external measurements point to a conclusion, it is time to look inside the machine. The Dolby Atmos Renderer 5.3.2 binary is a 71.7MB universal Mach-O protected by PACE Eden. The __text section is fully encrypted at rest. But the __cstring and __const sections are cleartext, and that’s where the truth is buried.

The binaural engine is statically linked. It consists of Dolby’s internal Fugu renderer library and the Rosella virtualizer. Embedded build paths point at /Users/builder/p4ws/renderer/fugu/build_external/. The third-party software document lists standard libraries: Xerces, asdcplib, BBC Audio Toolbox, CodeSynthesis XSD, FFTW, PACE iLok, Qt5, and Steinberg ASIO. There is no external HRTF library. No SOFA loader. No third-party binaural DSP. Everything spatial is in-house and proprietary.

C++ typeinfo reveals the processing graph: BinauralProcess::{AnalysisNode, SynthesisNode, ChannelSetupNode, RosellaMixerNode} and BinauralRenderer::{UpSamplerNode, DownSamplerNode, HeadphoneEqNode}. The actual processing happens in the RosellaMixerNode and the ProxyBinauralRenderer::PreprocessNode.

BinauralProcess
├── AnalysisNode
├── SynthesisNode
├── ChannelSetupNode
└── RosellaMixerNode            ← processing happens here
BinauralRenderer
├── UpSamplerNode
├── DownSamplerNode
└── HeadphoneEqNode
ProxyBinauralRenderer
└── PreprocessNode              ← and here
Exhibit AThe binaural processing graph, reconstructed from C++ typeinfo in Dolby Atmos Renderer 5.3.2.

The most damning evidence is in the __TEXT,__const section. There are four float32 coefficient blocks totaling ~4.6MB. The largest is 1,032,192 floats. These tables are spectrally smooth. The magnitudes cluster around 0.02-0.4. There are no impulse peaks and no headers. No impulse responses anywhere in 4.6MB: these are frequency-domain coefficients, magnitude spectra per direction and distance.

__TEXT,__const
  float32 coefficient blocks ........ 4
  total size ........................ ~4.6 MB
  largest block ..................... 1,032,192 floats
  magnitude cluster ................. 0.02 – 0.4
  impulse peaks ..................... 0
  headers ........................... 0
  phase information ................. none
Exhibit BThe coefficient tables. Smooth magnitude spectra per direction and distance. A magnitude spectrum cannot store a delay.

This storage format is a technical dead-end for anything that needs to model interaural delay. You cannot represent a delay in the frequency domain using just a magnitude spectrum. The phase information required for ITD is completely absent from these tables. This explains why the measurements show ITD tracking the panning floor. Dolby’s renderer is physically incapable of producing it.

The phase information required for ITD is completely absent from these tables.

The metadata confirms this. DBMD metadata carries a per-object binaural render mode: Off, Near, Mid, Far. Example metadata for a real ADM file shows: Ch10 Mid, Ch11 Mid, Ch12 Near, Ch13 Mid. These modes are room presets that control distance-dependent reverb. The Off mode gives the game away: no virtualization, with the channel still added to the binaural output at -3 dB. So by design, a portion of every binaural mix is just a downmixed pan. The LFE channel is always Off.

DBMD modeWhat it sounds like it doesWhat it does
Near · Mid · FarHRTF “binaural strength”Room preset for distance-dependent reverb
OffNothingNo virtualization. Channel still summed into the binaural output at −3 dB
LFEAlways Off

The Personalized Fantasy

In March 2022, Dolby announced personalized HRTF. The marketing copy was pure fantasy: “Even seemingly minor physical deviations can result in very different experiences.” The app would capture 50,000 points of the user’s head, ears, and shoulders. The promise was that Dolby would use this to customize the binaural experience to the individual user’s anatomy.

The binary reveals the pathetic reality of this feature. The Dolby Personalized Headphone Interchange (PHI) schema is embedded. A .personalized_headphone file’s personalized_hrtf payload contains exactly: rosella_coefficients (an int16 array from -32768 to 32767), rosella_coefficients_version '1.0.0', and a room_model field described as being incorporated in the coefficients.

.personalized_headphone
└── personalized_hrtf
    ├── rosella_coefficients           int16[]   -32768 … 32767
    ├── rosella_coefficients_version   '1.0.0'
    └── room_model                     (incorporated in the coefficients)
Exhibit CThe entire personalized payload. 50,000 capture points go in; a table of 16-bit coefficients for the same RosellaMixerNode comes out.

There is no separate personalization DSP. The typeinfo shows zero Personalized* node classes. Both the stock and personalized paths run through the exact same RosellaMixerNode. Personalization is simply a coefficient swap. When you load a personalized profile, the renderer swaps its stock magnitude tables for your personalized magnitude tables. It’s that simple.

The loader rejection strings confirm the crude nature of this implementation: Invalid pHRTF file: rosella serial tuning init failed. and Invalid pHRTF file: rosella serial tuning query memory failed. There is also a version-gate: Unsupported personalized_headphone file version. rosella_coefficients_version:.

And then there is the bundled DEE encoder, dee_ddpjoc_encoder. This 15MB unencrypted binary has 495 spatial_coding_* functions and exactly zero binaural, HRTF, or Rosella DSP functions. It only reads and writes the get/set_program_binaural_render_mode metadata using the DbmdYamlCodec::BINAURAL_RENDER_MODE key. The encode path does nothing but embed the word binaural as metadata. The actual binaural render happens in the downstream AC-4 immersive stereo decoder, not in Dolby’s production tools. This is why the v5.3.2 release notes mention Improved binaural headphone monitoring without mentioning personalization. The production tool never renders binaural at all; it writes the word into the bitstream and leaves the actual render to the consumer decoder.

dee_ddpjoc_encoder                      15 MB, unencrypted
  spatial_coding_* functions ........ 495
  binaural / HRTF / Rosella DSP ..... 0
  binaural touchpoint ............... DbmdYamlCodec::BINAURAL_RENDER_MODE
                                      (metadata key, nothing more)
Exhibit DThe encoder. The word “binaural” goes into the bitstream. No binaural processing does.

Personalized HRTF died of arithmetic. At a 15% HRTF blend, a 3-6 dB difference between generic and personalized HRTF manifests as a 0.5-1 dB difference at the output. The just-noticeable difference for level is ~1 dB. 50,000 capture points of a user’s head bought a sub-1 dB difference that sits below the threshold of audibility. That is the kind of number that gets a feature killed, and the feature died on cue: Dolby discontinued support to capture and download personalized profiles on July 1, 2025, one month after Apple announced ASAF and the APAC codec for visionOS. Apple’s ASAF is proprietary, with no Dolby support.

generic vs. personalized HRTF ......... 3 – 6 dB
× HRTF share of the output ............ 15 %
= difference at the ear ............... 0.5 – 1 dB
  just-noticeable difference, level ... ~1 dB
Exhibit EThe arithmetic that killed personalization.

Apple had already independently arrived at the same retreat. Analysis of the iOS 17 beta found that the 8-12 kHz high-frequency boost in Apple Spatial Audio was 90% completely eliminated. The feature sounded better with the filter turned down, so the filter got turned down.

The Naming Convention of Deception

Dolby’s naming history runs on this strategy. In the 1980s, Dolby Surround implied discrete multichannel but delivered matrix-decoded mono rear channels. In the 1990s, Dolby Digital implied perfect fidelity but delivered lossy compression. Atmos implied 3D immersion but delivered context-dependent delivery. Now, Binaural implies HRTF spatialization but delivers amplitude panning with reverb.

EraNameImpliedDelivered
1980sDolby SurroundDiscrete multichannelMatrix-decoded mono rear channels
1990sDolby DigitalPerfect fidelityLossy compression
Dolby Atmos3D immersionContext-dependent delivery
NowDolby “Binaural”HRTF spatializationAmplitude panning with reverb

The two lines of evidence converge perfectly. Output measurements show a system that matches panning in the time domain, decorrelates via reverb in the interaural difference domain, and provides minimal pinna-grade HRTF cues in the magnitude domain. Static analysis of the binary shows a system that uses frequency-domain coefficient tables, which cannot represent interaural delay, and implements personalization as a simple coefficient swap.

The HRTF is technically present and perceptually buried: panning sets the image, reverb fills the ears, and the pinna never gets a vote. The entire personalization pipeline was a marketing exercise in collecting data that yielded inaudible results. The binary contains the full confession: magnitude-domain processing, no ITD modeling, and a reliance on reverb for spatial decorrelation. Call it what the binary calls it: a panner with a reverb plugin and a very expensive marketing department.


The paper · PDFThe Emperor’s New Binaural: Reverse-Engineering Dolby Atmos Binaural Rendering Reveals Minimal HRTF ProcessingMethodology, CMAP derivation, full result tables, and limitations.