Next upHack for Humanity: San Francisco (powered by Google Gemini)
News

Together AI finds GLM-5.3 Flash 17x cheaper on DeepSWE

Together AI reports that GLM-5.3 Flash lagged the full model on first-attempt coding accuracy but cost roughly one-seventeenth as much per DeepSWE rollout. Its verifier-gated cascade reached 80.9% accuracy at $1.70 per task.

D
Sep 10, 2026 · 2 min read

Together AI says GLM-5.3 beat GLM-5.3 Flash on first-attempt coding accuracy in a 900-rollout DeepSWE comparison, while Flash cost about 17 times less per rollout. The August 28 report put pass@1 at 69.0% for GLM-5.3 and 63.4% for GLM-5.3 Flash.

Average rollout costs were $3.99 for GLM-5.3 and $0.24 for Flash, according to the publisher. Flash also averaged 26 minutes per run, against 35 minutes for the full model.

Together AI also tested a cascade: run Flash first, then escalate to GLM-5.3 if a verifier rejects the initial answer. The company reported 80.9% accuracy at $1.70 per task. By its figures, that put the cascade 11.9 percentage points above the full model’s single-shot pass@1 while cutting 57.4% from its $3.99 average rollout cost.

The comparison covered all 113 DeepSWE v1.1 tasks, with four trials planned for each model configuration at maximum effort. Together AI recorded 452 GLM-5.3 rollouts and 448 Flash rollouts. It reported one GLM-5.3 infrastructure error and four missing expected Flash trials, so some Flash tasks had only three trials. The company said its calculations used each task’s actual trial count.

Pass@1 is the average per-task success rate, with every task weighted equally. Pass@4 is the share of tasks solved in at least one of four attempts. Together AI said the full model’s 5.6-percentage-point lead at pass@1 shrank to 2.6 points at pass@4, with GLM-5.3 at 87.6% and Flash at 85.0%.

The official DeepSWE leaderboard separately shows rounded pass@1 scores of 69% and 63% and average costs of $3.99 and $0.24. It lists 124 average steps for GLM-5.3 and 123 for Flash; Together AI’s report gives 125 and 123. DeepSWE describes version 1.1 as 113 original, long-horizon engineering tasks across 91 active open-source repositories and five programming languages, each with a purpose-written functional verifier.

The findings add a cost comparison to DataPhoenix’s earlier coverage of GLM-5.3’s CyberGym benchmark. Together AI said Flash broke at least one already-passing baseline test in 6.9% of non-errored rollouts, versus 4.4% for the full model. The company reported that the two models together solved at least one trial on 102 of 113 tasks, or 90.3%.

The DeepSWE paper says every configuration uses a fixed harness. That makes the scores a comparison under common benchmark conditions, not a global ranking of coding products. Together AI said per-turn trajectory files were unavailable on the public CDN for both batches, leaving its analysis reliant on index-level records. Its report also omitted a complete verifier specification and false-acceptance and false-rejection rates for the cascade.

More news