Next upSF Pitch Night by the AI Collective - #SFTechWeek
News

Apple study finds LLM conditioning gains can reduce fluency

An Apple-published study found that methods for conditioning language models can improve control while degrading fluency, with activation steering behaving differently in base and instruction-tuned models.

D
Sep 30, 2026 · 2 min read

An Apple-published study found that methods used to condition large language models can improve control over generated content while reducing fluency and instruction following.

The practical consequence, according to the study, is that a method can appear successful when evaluated only on whether it injects or removes a target concept, even as its output becomes less natural or responsive. The authors also found materially different behavior between base and instruction-tuned models: activation steering barely exceeded the unconditioned baseline for concept injection on instruction-tuned models. Those models often ignored the target concept while retaining stronger fluency and instruction following.

The underlying paper compared prompting, LoRA-based supervised fine-tuning and three activation-steering methods: ITI, CAA and Linear AcT. The researchers tested concept injection and concept removal across base and instruction-tuned versions of Qwen3-0.6B, Qwen3-8B, Qwen3-14B and SmolLM3-3B.

The paper is separate from Apple’s recent study of losses when language models serialize structured reasoning. That research examined information preservation during text serialization, not conditioning methods.

For concept injection, the conditioning study covered 60 concepts in 10 categories. The researchers generated 400 training sentences per concept and, for each combination of concept, method and model, produced 1,000 continuations of up to 100 tokens from a pool of 100 generic prompts.

Concept removal had a narrower scope. The experiment focused only on toxicity mitigation, using training subsets derived from RealToxicityPrompts and Thoroughly Engineered Toxicity prompts for generation.

The authors reported that prompting and supervised fine-tuning outperformed activation steering on both concept relevance and fluency in the injection tests. In the toxicity-removal experiment, however, all three activation-steering methods were more effective than prompting and supervised fine-tuning. When activation steering achieved strong concept injection, the study associated those gains with lower fluency and instruction following, partly because continuations repeated text.

The paper also tested evaluation measures against an OLMo2-32B-IT judge. The authors reported Spearman correlations of 0.97 between Concept Similarity and the judge’s concept-injection score, 0.96 between ToxClass and concept removal, and 0.83 between Instruction Similarity and instruction following.

Perplexity had a correlation of zero with the judge’s fluency score, compared with 0.81 for type-token ratio. The authors argued that predictable repetition can make poor text appear deceptively fluent under perplexity.

The researchers cautioned that the study covered a relatively small set of models, excluded very large models and examined concept removal only through toxicity mitigation. The opened sources do not include an independent replication, and the paper’s human-validation annotations were performed by its authors.

More news