Next upSF Pitch Night by the AI Collective - #SFTechWeek
News

Apple study finds independent token sampling can distort distributions

Apple researchers found that confidence-ordered discrete diffusion can pass per-sample checks yet miss the target distribution when dependent token positions are sampled independently.

D
Oct 3, 2026 · 3 min read

Apple researchers found that independent-update discrete diffusion samplers can miss a training distribution even when every generated sample passes task-level checks. In a synthetic task, the Apple study found that a confidence-ordered sampler produced 2,000 outputs with 100% well-formedness, correctness and uniqueness, while its pooled-token distribution remained far from the target.

At a temperature of 0.5 and 30 sampling steps, the study measured pooled-token total variation of about 0.157, roughly 29 times a 0.0054 sampling-noise floor. The comparison highlights what the two tests measure differently: per-sample metrics assess outputs one at a time, while distribution-level accuracy tests whether outputs occur with the right frequencies.

The work, by Russ Webb, Amitis Shidani, Alice Bizeul and Dan Busbridge, examines samplers that update a group of token positions at once. Each position is drawn independently from its own predicted distribution, so the update’s joint distribution is a product of those per-position distributions. It adds to Apple research on structure-sensitive model behavior, including a study measuring losses when language models serialize structured reasoning.

The paper’s Theorem 3 states that this product cannot equal the group’s true conditional distribution when the positions remain dependent after the other tokens are fixed. The divergence is bounded below by the group’s total correlation, even if every per-position prediction is exact. Theorem 4 constructs joint distributions with identical per-position marginals but different dependencies, showing that a scheduler using only those marginals cannot certify that a nontrivial group is conditionally independent.

The authors conclude that an independent-update sampler can reproduce the training distribution only when every group of two or more positions updated together is conditionally independent given the frozen tokens. The certification result does not mean every grouped update is wrong. It applies to schedulers that see only per-position distributions; a scheduler with additional information about the data distribution falls outside the theorem’s assumption. Product-form updates must still avoid dependent groups.

The experiments used ScanAndAdd, a synthetic 25-token task in which an answer is determined by nine values and a block of 10 commands containing two add operations. In the reported M=1 condition, the target was uniform over 45 billion valid sequences and had entropy of 35.39 bits. The model was a six-layer bidirectional Transformer with 4.75 million trainable parameters, trained for 60,000 steps on 9 million examples. Each reported measurement used 2,000 generated samples.

Those perfect scores did not cover every constraint in the target distribution. The well-formedness and correctness metrics did not verify that a generated command block contained exactly two add commands. In two manual orders that updated all 10 command positions in one step, 70.3% and 70.7% of samples had the wrong command count, close to the study’s 69.8% analytical prediction. Those samples had zero probability under the target distribution while still scoring as well-formed and correct.

The same trained model produced a different measured outcome when the ordering changed. Hand-specified schedules that sampled independent values in parallel, commands sequentially and answers last reached pooled-token total variation between 0.0050 and 0.0066, with correctness between 98.7% and 99.6%. At the available sample size, the paper’s tests did not distinguish those token distributions from fresh draws from the training data.

The measured distortion comes from one synthetic task, not a natural-language, image or speech corpus. The authors did not train a uniform-state diffusion model, and the Section 5 measurements were limited to absorbing-state or remasking samplers in the M=1 condition.

More news