Next upHack for Humanity: San Francisco (powered by Google Gemini)
News

GenCeption turns a video-diffusion model into a general-purpose vision learner, DeepMind team says

Google DeepMind researchers introduced GenCeption, a model that repurposes a video-diffusion backbone to match specialist vision models using up to 500 times less training data.

Dmytro Spodarets
Jul 13, 2026 · 1 min read

Google DeepMind researchers have introduced GenCeption, a vision model that repurposes a pretrained text-to-video diffusion backbone to handle tasks like depth estimation and segmentation, in a paper posted to arXiv on July 10.

The claim of note is efficiency: the authors report state-of-the-art or matching results against specialized models while using 7 to 500 times less training data, suggesting a single generative backbone can absorb skills that today require task-specific systems.

GenCeption adapts the diffusion model into one feed-forward network steered by text instructions, with reported results on depth and surface-normal estimation, camera-pose estimation, referring segmentation and 3D keypoint prediction. It is trained mostly on synthetic video, including synthetic human footage, yet generalizes to real-world video and to unusual categories such as animals and robots, the authors say. The work, led by DeepMind’s Letian Wang with co-authors including Andrew Zisserman, Joao Carreira and Kaiming He, has been accepted at ECCV 2026.

The results come from the authors’ own paper and have not been independently reproduced. The heavy reliance on synthetic training data means real-world robustness across domains remains to be tested at scale.


Dmytro Spodarets
Dmytro Spodarets
Founder & Editor-in-Chief

Founder and Chief Editor of Data Phoenix — a San Francisco Bay Area media and education platform focused on AI and Data.

More news