Together AI adds canary rollouts for production model upgrades
Together AI added staged rollout controls for Dedicated Model Inference, including health checks and optional metric gates, but failed gates pause for operator review rather than automatically restoring the old model.
Together AI introduced canary rollouts for its Dedicated Model Inference service, giving operators a staged way to move production traffic from one model deployment to another on the same endpoint. The controls add health checks and optional metric gates before a new deployment receives all requests, while leaving the endpoint URL unchanged.
The rollout control serves a different purpose from Together AI’s endpoint-level A/B testing for production models: the rollout feature governs a migration between deployments, while A/B testing compares live model variants.
The documented rollout process moves traffic from a source deployment to a target deployment in steps. By default, the canary sequence sends 5%, 25%, 50% and then 100% of traffic to the target, with a soak period after each step. Custom percentages are also supported.
At every canary step, Together AI says the platform scales the target first, checks its health, shifts traffic, waits for routing changes to propagate, drains corresponding source capacity, and then applies the soak period and any configured metric gate. The step is recorded as complete only after those checks pass.
Metric gates are optional and available only for canary rollouts. According to the metric-gate documentation, operators can evaluate router latency, router error rate or in-flight requests against either a fixed threshold or the source deployment’s result. Every configured rule must pass before the rollout advances. Blue-green and rolling strategies do not accept metric gates because they do not have the required soak window.
The same documentation describes an edge case: a fully idle endpoint skips the metric gate and proceeds without it. If the endpoint has traffic but the target produces no samples, the metric is marked unavailable and the rollout pauses for review.
A failed gate does not automatically return all traffic to the old model. It puts the rollout in SYSTEM_PAUSED at the current traffic split. An operator can resume and re-evaluate the gate, promote the target, or cancel the rollout. Canceling freezes the current weights and leaves both deployments serving those shares. The rollout guide says there is no separate rollback operation; returning traffic requires a new rollout with the source and target reversed.
That recovery path is narrower than the announcement’s “automatic rollback” description. The documented automation is the pause on a failed gate, while cancellation and the reverse rollout remain operator actions.
Through the REST API, creating a rollout leaves it in a pending state until a separate start request. Together AI’s CLI combines creation and startup in one command. While a rollout is active, including while paused, the platform blocks traffic-split changes, permits only one active rollout per endpoint, and prevents either participating deployment from being stopped or deleted.
Together AI reported a company-run demonstration in which 10% of traffic moved from Qwen2.5-7B-Instruct to Qwen3.5-9B. The company said the rollout paused after measured p95 latency rose from about 734 milliseconds to 1,740 milliseconds, then was canceled at the 90/10 split and reversed. It also said 6,800 requests were served without a non-200 response during the demonstration. Those results have not been independently verified and do not establish that every rollout will avoid interruption.
The rollout documentation is inconsistent about a stopped target whose minimum and maximum replica counts are both zero: one requirement says a stopped target is eligible and will restart, while startup guidance warns that this setup can move traffic before a replica is ready and cause deployment_stopped errors. The opened materials do not state whether the rollout capability is generally available or covered by a service-level commitment; the documented CLI and SDK interfaces are labeled beta.
More news

xAI launches Team Bots for shared workflows

AWS adds xAI’s Grok 4.7 to Amazon Bedrock

OpenAI adds $5 million and up to $5 million in credits to Lenfest AI program
