Next upPhysical AI VC <> Founders Pitch Night #SFTechWeek @Mission Robotics
News

Together AI says IBM Cloud B300 inference cluster is running

Together AI says it is running open-model inference on a dedicated NVIDIA B300 cluster hosted by IBM Cloud, moving a $240 million infrastructure agreement from planned capacity to a vendor-reported operational milestone.

D
Oct 7, 2026 · 2 min read

Together AI said it is now the first customer running on a dedicated, large-scale NVIDIA B300 inference cluster on IBM Cloud. The October 6 update moves the project from an announced capacity plan to a vendor-reported operational milestone. IBM and NVIDIA have not independently confirmed in the supplied evidence that the cluster is running.

It remains unclear whether the live capacity is the full contracted cluster, an initial tranche or capacity operating before broader customer availability.

IBM announced the multi-year, $240 million agreement on August 11. IBM said it would deploy its first dedicated, large-scale inference cluster built with NVIDIA HGX B300 systems and Spectrum-X Ethernet for Together AI, with availability expected in the first quarter of 2027. Together AI’s new post does not explain how the reported milestone affects that timetable.

Together AI describes the current setup as a dedicated B300 GPU cluster purpose-built for inference and backed by Spectrum-X Ethernet networking. The companies say Together AI operates the inference layer, IBM supplies the cloud infrastructure, and NVIDIA provides the GPUs and networking. For related context on the company’s production tooling, see its staged rollout controls for Dedicated Model Inference.

The research sources do not disclose the number of GPUs, HGX systems, nodes or racks, the physical region, the exact commissioning date, current utilization or customer-access terms.

Together AI also says its platform processes hundreds of trillions of tokens per month for more than one million developers, with demand rising. IBM’s August release cited 400 trillion tokens per month. Neither company provided independent usage evidence in the supplied sources. Together AI’s claims about performance, cost, reliability, security and guardrails are also unsupported there by comparative benchmarks, pricing, workload data, service-level results or customer evidence.

NVIDIA’s HGX platform specifications list eight Blackwell Ultra SXM GPUs and 2.1 TB of total memory for one HGX B300 baseboard. Those per-system specifications do not establish how many systems or GPUs IBM Cloud has deployed for Together AI.

More news