Reproducible audio codec benchmark · 2026

DACOPUSMP3

How much fidelity fits in an hour? A 170-clip study of storage, perceptual degradation, spectral error, intelligibility and neural decoding speed.

Packed DAC bitrate7.76kbps measured
Storage per audio hour3.49MB / hour
L40S decode speed124×realtime median
DAC decoder network54.1Mparameters
01 · Storage economics

Small tokens,
large differences.

Packed all-codebook DAC is the smallest representation in this study. Opus 16 kbps uses 2.16× its storage; MP3 128 kbps uses 16.61×. Costs below use marginal R2 Standard storage before free tier or billing rounding.

Measured payload and container sizes, normalized to one hour. Hover a bar for measured bitrate, size, monthly R2 cost and relative storage.
RepresentationkbpsMB/hour$/hour-month
DAC packed tokens7.763.49$0.000052
DAC uint16 tokens12.415.58$0.000084
DAC native file12.835.77$0.000087
Opus 16 CBR16.757.54$0.000113
Opus 48 CBR48.8021.96$0.000329
Opus 64 CBR64.8229.17$0.000438
MP3 128 CBR128.8257.97$0.000870
02 · Objective quality

One score is
never enough.

Twenty naturally challenging LibriSpeech utterances and twenty matched speech-plus-music mixtures expose different failure modes. We report waveform fidelity, intelligibility, multi-resolution spectral distance and a learned MOS estimate separately.

Mean with 95% clip-level bootstrap intervals. ↑ higher is better for SI-SDR and STOI; ↓ lower is better for log-STFT distance and MOS loss. Use the legend to isolate a cohort.
Key result

DAC beats Opus 16

At roughly half the storage, packed DAC has better SI-SDR and substantially lower spectral distance in both stress cohorts.

Speech

Intelligibility holds

STOI stays high for every codec. Higher-rate Opus and MP3 approach one, while DAC remains strong at its much smaller payload.

Caution

MOS is secondary

SQUIM is trained for speech and can behave oddly on music mixtures. Near-zero or negative deltas are model-level ties, not proof of improvement.

03 · Neural decoding

Realtime,
comfortably.

The 54.104M-parameter synthesis decoder runs at 2.68× realtime on this 12-thread CPU and 124.38× on an NVIDIA L40S. Including quantizer lookup, the complete decode path is 54.344M parameters.

Warm-cache, batch-one application-level decode of 8.29 seconds. Hover for ranges and benchmark details.
Benchmark boundary

What the timer includes

Codebook reconstruction, neural synthesis, GPU-to-CPU output materialization, loudness normalization and output-length restoration. Disk I/O, model loading and encoding are excluded.

CPU: 2.65–2.74× realtime
GPU: 111.0–139.6× realtime
Full codec: 76.652M parameters

04 · Measurement

Metrics with
clear jobs.

Codec evaluation is multi-dimensional. The suite follows common neural codec practice while documenting where a metric is uncalibrated, speech-specific or unavailable in a safe reproducible build.

Included

SI-SDR ↑

Strict, scale-invariant waveform reconstruction fidelity. Phase and timing changes count even when they sound benign.

Included

STOI ↑

Short-time speech intelligibility. Applied only as a speech-component diagnostic.

Included

Log-STFT L1 ↓

Three spectral resolutions capture both envelope and fine-structure mismatch.

Included

SQUIM MOS loss ↓

Learned speech-quality prediction, paired against each pristine reference.

Researched

ViSQOL

Common full-reference codec metric. Not scored here because no safe official wheel supports this Python build.

Follow-up

MUSHRA

A controlled human listening study remains the gold standard for final perceptual decisions.

20 examples · 6 conditions each

Hear the tradeoffs.

Switch between the reference, DAC, three Opus rates and MP3 while keeping playback position aligned.

Open listening lab →
05 · Provenance

Built to be
reproduced.

LibriSpeech test-other was checksum-verified and sampled deterministically across 20 speakers. It is naturally challenging read speech—not artificially noisy speech. MTG-Jamendo tracks come from an artist-disjoint test split with adaptation-compatible licenses, mixed at +5 dB active-speech SNR.

All exact hashes, offsets, gains, timings and raw measurements are downloadable below.