English

Model ReleasesApple

Apple Proposes Latent-Space Distillation to Compress Streaming Neural Audio Encoders

Apple researchers have introduced a method to compress streaming neural audio encoders using latent-space distillation. The technique aims to reduce the memory footprint of the audio tokenizer—the component responsible for mapping waveform windows into representations for language models—to alleviate memory competition with the foundation model on-device.

Unlike traditional distillation that targets discrete tokens or output distributions, this approach uses the pre-quantizer latent representation as the supervision target. This allows the distilled student model to cover multiple token interfaces and benefit from models pre-trained either alone or jointly with a language model.

Experimental results show that at 2.8× compression, the distilled student model maintains a relative Word Error Rate (WER) within 1.9% of the teacher model across five of six tested pairs without fine-tuning. The method also improved performance by 3.9% relative compared to an independently trained tokenizer of identical capacity.

Sources

  1. Compressing Streaming Neural Audio Encoders via Latent-Space Distillation (Apple Machine Learning, 2026-09-24)
  2. arXiv