Level-matched A/B: building an honest comparison
Louder tends to win a quick comparison whether or not it deserves to. Why an A/B instrument should make loudness bias explicit, how the Studio's player does it, and what a level-matched switch can and cannot tell you.
Learned building WPAudio Engine
Written and engineered by WPAgency Studio
The mastering case study on this site ends with a listening instrument: twelve tracks, each in two versions — original pre-master and reference master — behind one transport. The obvious way to use it is to switch back and forth and decide which you prefer. The problem is that the reference masters measure between 3.1 and 6.5 LU louder than the originals, and that difference alone is enough to tilt the answer.
This article is about that tilt: where it comes from, what the instrument does about it, and — just as important — what controlling it does not prove. As in the loudness article, general knowledge and this Studio’s own evidence are kept explicitly separate.
The problem
Switch quickly between two versions of the same material and the louder one tends to come across as fuller, clearer, more present — even when the only difference is gain. The pattern and its consequences for mastering practice are discussed in Earl Vickers’ “The Loudness War: Background, Speculation, and Recommendations” (AES 129th Convention, 2010), the source this Studio already leans on for the history of escalating release loudness: when louder reliably reads as better, every uncontrolled comparison quietly rewards whoever pushed the level hardest.
For anyone evaluating a master — including the engineer who made it — that is a trap. A pre-master versus reference-master switch is supposed to answer “what did the mastering do?” If the two versions differ by several decibels of loudness, the honest answer is: mostly, it got louder, and your ears will report that as quality whether or not anything else improved.
Why level matching matters
The fix is old and simple: compare at matched loudness. Attenuate the louder version until both sit at the same measured level, then switch. Whatever preference survives the match is about the content of the processing — the balance, the density, the ceiling behaviour. Whatever preference does not survive it was about the gain.
One thing level matching is not: a cure for bias in general. It controls one well-documented variable. A sighted comparison is still sighted, expectation still leans on the label, and short passages still get judged differently from whole songs. Matching the level just removes the loudest of the thumbs from the scale.
How the Studio’s instrument approaches it
What follows describes the shipped behaviour of the A/B player, all of it public and verifiable in a browser:
- Each track loads both versions — the original 16-bit WAV pre-master and the 24-bit reference master, the actual production files. The master is published in two lossless containers, FLAC and ALAC, because no single one decodes in every browser; both are encoded from the same audio and the build verifies that they produce identical samples, so the comparison does not depend on which container a listener’s browser accepts.
- Switching versions keeps the playhead, so an A/B is a change of master, not a restart of the song.
- Nothing autoplays and nothing preloads until the listener acts; the comparison is entirely user-driven.
- The transport is keyboard-operable, and loading and failure states are announced rather than silent.
- Level matching is an explicit, labelled control — off by default, stating the exact attenuation it will apply, with a note that some mobile browsers fix media volume and the toggle has no effect there.
Off by default is deliberate. The unmatched comparison is real too — the masters genuinely are louder, and a listener should hear that fact. The toggle exists so that the reason for a preference can be separated from the preference itself, and so the loudness difference is stated as a number rather than left to masquerade as quality.
The measurement
The attenuation is not a house guess; it is each track’s measured loudness difference. Both versions of every track were measured to ITU-R BS.1770 integrated loudness — the published values live in the case study — and the player attenuates the louder version by the difference, converting decibels to a linear gain in the standard way:
gain = 10^(−Δ/20)
Track 01 is the worked example: the pre-master measures −16.1 LUFS and the reference master −10.7 LUFS, a delta of 5.4 LU. The instrument therefore plays the master at 10^(−5.4/20) ≈ 0.537 of full volume — the label reads “−5.4 dB on the master” — bringing it to the original’s integrated loudness. Every track uses its own measured delta; across the album those deltas run from +3.1 to +6.5 LU.
One honest limitation, stated rather than hidden: matching integrated loudness aligns the whole-track average. It does not make every moment of the two versions equally loud — a master with different dynamics will still be momentarily louder or quieter than the original in places. Integrated matching is the right single correction, not a perfect one.
What this does not prove
Level matching does not tell anyone which version is artistically better. It does not replace critical listening, and it does not replace the engineering judgment the mastering itself required. It cannot prove that one version is objectively superior — no measurement can, because “better” is not a defined quantity. And this Studio has published no listening-test results: the instrument collects no preferences and we make no claim about what listeners choose. The claim is narrower and checkable: with the toggle on, the loudness variable is controlled to the measured values published beside the audio.
What Wings & Prayers demonstrated
The album gave the principle its production test. The deltas are large enough to matter — up to 6.5 LU on track 09 — which means an unmatched switch would have favoured every reference master by gain alone, on every track, before a single mastering decision got a hearing. Building the match into the instrument was the difference between publishing a demo and publishing a comparison. That decision, and the rule behind it — publish only what you can measure — is recorded in the journal entry and in ADR 0007.
Practical takeaway
Before concluding that processing sounds better, make sure you are not simply responding to a level difference. Measure both versions, match the louder to the quieter, and only then ask your ears the question. If the preference survives the match, it is about the work. If it does not, it was about the volume knob — and now you know.
References
- E. Vickers, The Loudness War: Background, Speculation, and Recommendations, AES 129th Convention, 2010 — AES E-Library, paper 15598
- ITU-R Recommendation BS.1770-5, Algorithms to measure audio programme loudness and true-peak audio level, 11/2023 — itu.int/rec/R-REC-BS.1770-5-202311-I
- The measured per-track values used by the instrument are published in the mastering case study; the three numbers themselves are explained in How to read a loudness report.