DAC beats Opus 16
At roughly half the storage, packed DAC has better SI-SDR and substantially lower spectral distance in both stress cohorts.
How much fidelity fits in an hour? A 170-clip study of storage, perceptual degradation, spectral error, intelligibility and neural decoding speed.
Packed all-codebook DAC is the smallest representation in this study. Opus 16 kbps uses 2.16× its storage; MP3 128 kbps uses 16.61×. Costs below use marginal R2 Standard storage before free tier or billing rounding.
| Representation | kbps | MB/hour | $/hour-month |
|---|---|---|---|
| DAC packed tokens | 7.76 | 3.49 | $0.000052 |
| DAC uint16 tokens | 12.41 | 5.58 | $0.000084 |
| DAC native file | 12.83 | 5.77 | $0.000087 |
| Opus 16 CBR | 16.75 | 7.54 | $0.000113 |
| Opus 48 CBR | 48.80 | 21.96 | $0.000329 |
| Opus 64 CBR | 64.82 | 29.17 | $0.000438 |
| MP3 128 CBR | 128.82 | 57.97 | $0.000870 |
Twenty naturally challenging LibriSpeech utterances and twenty matched speech-plus-music mixtures expose different failure modes. We report waveform fidelity, intelligibility, multi-resolution spectral distance and a learned MOS estimate separately.
At roughly half the storage, packed DAC has better SI-SDR and substantially lower spectral distance in both stress cohorts.
STOI stays high for every codec. Higher-rate Opus and MP3 approach one, while DAC remains strong at its much smaller payload.
SQUIM is trained for speech and can behave oddly on music mixtures. Near-zero or negative deltas are model-level ties, not proof of improvement.
The 54.104M-parameter synthesis decoder runs at 2.68× realtime on this 12-thread CPU and 124.38× on an NVIDIA L40S. Including quantizer lookup, the complete decode path is 54.344M parameters.
Codebook reconstruction, neural synthesis, GPU-to-CPU output materialization, loudness normalization and output-length restoration. Disk I/O, model loading and encoding are excluded.
CPU: 2.65–2.74× realtime
GPU: 111.0–139.6× realtime
Full codec: 76.652M parameters
Codec evaluation is multi-dimensional. The suite follows common neural codec practice while documenting where a metric is uncalibrated, speech-specific or unavailable in a safe reproducible build.
Strict, scale-invariant waveform reconstruction fidelity. Phase and timing changes count even when they sound benign.
Short-time speech intelligibility. Applied only as a speech-component diagnostic.
Three spectral resolutions capture both envelope and fine-structure mismatch.
Learned speech-quality prediction, paired against each pristine reference.
Common full-reference codec metric. Not scored here because no safe official wheel supports this Python build.
A controlled human listening study remains the gold standard for final perceptual decisions.
Switch between the reference, DAC, three Opus rates and MP3 while keeping playback position aligned.
LibriSpeech test-other was checksum-verified and sampled deterministically across 20 speakers. It is naturally challenging read speech—not artificially noisy speech. MTG-Jamendo tracks come from an artist-disjoint test split with adaptation-compatible licenses, mixed at +5 dB active-speech SNR.
All exact hashes, offsets, gains, timings and raw measurements are downloadable below.