Apple researchers shrink streaming audio encoders with latent-space distillation
Apple researchers report compressing streaming audio encoders by 2.8-fold or more while five of six tested students stayed within 1.9% relative word error rate of their teachers.
Apple researchers used latent-space distillation to shrink streaming neural audio encoders by 2.8-fold or more, reporting that five of six tested student models stayed within 1.9% relative word error rate of their teachers without fine-tuning.
The research targets the audio tokenizer used as the front end for system-wide Dictation, which Apple says runs entirely on-device. The tokenizer processes every audio frame and shares a fixed memory and compute budget with the foundation model it feeds. A smaller encoder can therefore leave more of that budget available to the language model. The paper reports reductions in parameters and counted per-frame compute, not direct measurements of device RAM, latency, energy use or battery life.
An audio tokenizer turns waveform segments into representations a language model can consume. In the reported system, 16 kHz audio becomes one latent vector every 80 milliseconds. The term here refers to an audio front end; separate Apple research on protein-model tokenization addresses a different technical problem.
Rather than train a smaller encoder to copy discrete token assignments, the researchers froze the larger teacher and trained only the student to match its per-frame representation immediately before quantization. A single affine layer handled the width difference between teacher and student, and the same target worked for both discrete-token and continuous interfaces.
The paper evaluated six teacher-student pairs across two training stages. The three students distilled before joint language-model training recorded relative average-WER changes of 0.4%, 1.9% and 0.8% worse than their teachers. Results after joint training varied more: one student was 1.5% worse, one was 5.3% better and one was 7.7% worse. The last result was the only pair outside the paper’s five-of-six summary.
Word error rate was measured through full recognition systems, using a fixed 262-million-parameter transcription decoder for the first-stage pairs and a fixed 2.8-billion-parameter joint language model for the second-stage pairs. The two stages used different public benchmark suites, which the authors said should not be compared directly.
At equal student capacity, the distilled first-stage model averaged 11.90 WER, compared with 12.36 for an independently trained tokenizer, a reported 3.9% relative improvement. The authors noted that this control inherited four initialized sub-blocks from a distilled model, so it was not a clean from-scratch baseline.
The authors said the production configuration is tuned separately and differs from the research setup. The reported WER figures are public-benchmark results under the paper’s research protocol, not measured production-device performance or a product guarantee.
More news

xAI launches Team Bots for shared workflows

AWS adds xAI’s Grok 4.7 to Amazon Bedrock

OpenAI adds $5 million and up to $5 million in credits to Lenfest AI program
