MBZUAI researchers propose parallel decoding to speed video object tracking
MBZUAI says Parallel Tube Decoding generates a video’s object-location boxes simultaneously, reducing sequential steps and limiting how one bad box can distort later predictions.
Researchers at Mohamed bin Zayed University of Artificial Intelligence and the University of California, Merced have developed Parallel Tube Decoding, a method that identifies when a queried event occurs in a video and then generates the object’s location boxes in parallel instead of building its trajectory one box at a time.
In a controlled VidSTG comparison, the researchers reported tube-completion latency of 0.4 seconds and throughput of 45.9 boxes per second, versus 31.6 seconds and 0.5 boxes per second for an unquantized autoregressive baseline. They characterized those results as 79-fold lower latency and 92-fold higher throughput. The released model card says the measurements used BF16, batch size 1, a single 64-GB AMD Instinct MI210 GPU and synchronized decode-only timing. The results have not been independently reproduced.
The method addresses spatio-temporal video grounding, in which a model must find the interval containing a described event and locate the referred person or object throughout it. Conventional generative systems serialize that trajectory, making each new box dependent on earlier output. The research paper describes PTD as a two-stage process: it predicts one temporal block for the event interval, then generates all time-conditioned spatial blocks for that interval at once. This fixes sequential decoding depth at two rounds regardless of the number of boxes, compared with 1 + T rounds for sequential block decoding, where T is the number of boxes.
PTD uses a mechanism called Decoupled Block Attention. Each spatial block can attend to the shared video and text query and to the predicted temporal interval, but not to other spatial blocks. Each box is therefore grounded against the source video rather than against previously generated boxes, allowing the boxes to be computed together.
The researchers also linked sequential generation to compounding localization errors. In their history-correction experiment, replacing one mislocalized box with the ground-truth box improved later boxes, and the benefit faded as the distance from the corrected box increased. Their attention analysis found that sequential block decoding shifted attention from the video toward earlier localization text as a trajectory grew, while PTD maintained attention to the video. These are author-reported experimental findings, not an independent validation of the proposed mechanism.
Longer trajectories had little effect on PTD latency in the researchers’ tests: increasing a tube from eight to 64 boxes raised latency from 0.33 to 0.40 seconds, while sequential block decoding increased from 0.72 to 6.18 seconds. The team also reported that its GRPO-trained, 4-billion-parameter Qwen3-VL model led seven of eight VidSTG metrics and improved mean video intersection-over-union over the prior best in its comparison by 1.1 points for both declarative and interrogative queries.
For zero-shot referring video object tracking, PTD generates bounding-box prompts and passes them to off-the-shelf SAM2 or SAM3 segmentation systems. The reported tracking comparisons therefore do not represent a standalone PTD-only segmentation pipeline. The project’s results cover Ref-DAVIS, Ref-YT-VOS and ReasonVOS.
The project is separate from MBZUAI’s recent research on identity-safe image editing, but both concern computer-vision systems that operate on visual media.
The authors released a Qwen3-VL-4B-Instruct checkpoint and inference code, but not the training data, GRPO data or GRPO data-selection code. The current formulation predicts one continuous time interval and one spatial tube per query. It is not designed for disjoint event occurrences or multiple simultaneously valid objects, and the researchers list small, fast-moving or occluded targets and subtle state changes among its limitations.
More news

Mistral launches Large 4 in public preview

TII launches Falcon-Emirati-7B for Emirati Arabic

Google Research lays out unresolved privacy and security risks for AI agents
