Google Research introduces Diffusion Controller to steer image generation
Google Research introduced Diffusion Controller, a lightweight side network that steers prompt alignment while a diffusion model stays frozen in its gray-box configuration.
Google Research introduced Diffusion Controller, a framework that uses a lightweight side network to steer image generation. In the gray-box configuration, developers with access to intermediate denoising outputs can train the correction network while leaving the pretrained model’s weights frozen, then adjust the strength of that control during inference.
The authors cast reverse diffusion—the step-by-step process that turns noise into an image—as a stochastic control problem. The paper says the controller reweights the pretrained model’s reverse-time transitions toward a target while an f-divergence cost limits deviation from the original process. Its score function combines a fixed baseline from the pretrained model with a learned correction conditioned on intermediate denoising output.
The authors describe different training routes for different kinds of feedback. Supervised fine-tuning can be used when paired targets are available. If training provides only a final reward score, the framework supports reward-weighted regression or policy-gradient updates, including a PPO-style rule. Google says a single guidance-strength parameter can increase or reduce the controller’s influence at runtime, changing the balance between prompt alignment and the baseline output.
Google tested the method with a Stable Diffusion v1.4 backbone across supervised fine-tuning, reward-weighted loss and PPO. The company says the frozen gray-box controller beat LoRA on HPS-v2 win rate in the supervised and reward-weighted tracks. Those results are author-reported and were not independently reproduced in the reviewed sources. HPS v2’s authors describe it as a learned preference-prediction model for comparing images generated from the same prompt, so its win rates are automated metric results rather than direct human votes.
Google separately reported that a fully unlocked white-box version achieved a 90% win rate over the pretrained baseline. That figure does not apply to the frozen gray-box setup: Google’s white-box Diffusion Controller-J and Diffusion Controller-S variants also train the base model. The blog does not specify alongside the 90% claim which training regime, prompt count, checkpoint or uncertainty interval produced it.
DataPhoenix has separately covered Google Research’s multi-agent frameworks for coherent long-form video.
Google Research lists the work for ICML 2026, and the paper was submitted to arXiv on March 7. The reviewed sources do not establish that the method works through a fully opaque image-generation API; the technical description of gray-box adaptation requires exposed intermediate denoising outputs.
More news

EliseAI raises $350 million at a $4 billion valuation

California nonprofit sues OpenAI over Hugging Face intrusion

Trump orders federal agencies to use ‘Super Intelligence’
