Next upAI x Bio Pitch Contest
News

Together AI adds live tracking and finer controls to fine-tuning

Together AI expanded its fine-tuning service with live metrics, dataset inspection, in-flight run comparisons and more granular training controls, along with support for additional open-weight models and selected price cuts.

D
Sep 26, 2026 · 3 min read

Together AI has expanded Together Fine-Tuning, its model customization service, with live experiment tracking, dataset inspection, side-by-side run comparisons and more granular training controls. Announced September 11, the update gives developers more visibility before and during training, more control over how data shapes a run, and a way to stop jobs that are no longer improving.

The service now records metrics at every training and evaluation step. Developers can access them through the dashboard, command-line interface, API and software development kit, according to Together AI’s monitoring documentation. Training series include loss, gradient norm and learning rate. The dashboard can also chart multiple selected jobs while they are running, using aligned sampling and controls for comparing a chosen range of steps. When a run includes a validation file and evaluations are enabled, validation loss and other evaluation metrics appear alongside the training data.

New early stopping controls use validation loss to decide when a run should end. Developers can set patience, a minimum improvement threshold and a warmup period. If validation loss stops improving within those settings, Together AI says the job finishes with the best checkpoint promoted and unused training steps refunded.

Dataset inspection now covers both ends of a run. Before training begins, developers can preview tokenized rows, including token IDs, token strings, labels, trained spans and any truncation. Once a job has produced its tokenized dataset archive, developers can download it from the job details. Uploaded files also go through server-side schema validation during ingestion, and data that fails validation is marked INVALID_FORMAT with an error before it can be used for training.

Developers also get finer control over training data and execution. Root-level sample weights can scale an entire JSONL example’s contribution to training loss. Sequence packing remains on by default for JSONL data but can be disabled, while pre-tokenized Parquet files can include custom attention masks, labels or position IDs. Gradient accumulation supports an effective batch size larger than a single micro-batch; Together AI defines that effective size as the batch size multiplied by the number of accumulation steps.

The supported-model list now includes announced additions such as GLM 5.3, DeepSeek V4 Flash variants, Qwen3.8 27B, Kimi K2.7 Code and Gemma 4 models for LoRA fine-tuning. LoRA, or low-rank adaptation, trains small adapter weights rather than updating every parameter in the base model. Together AI also says adapters can now be attached to expert modules in mixture-of-experts models. In a company-run test using 200 invented facts, Together AI reported recall of up to 89% for expert-layer adapters, versus 15% for attention-only adapters. The supplied research did not independently reproduce that result.

Together AI says it cut training prices for most supported models by 30% to 70%. Its current pricing page lists LoRA training for Qwen3.5 9B at $0.34 per million tokens for supervised fine-tuning and $0.84 for direct preference optimization. For gpt-oss-20B, the listed prices are $0.40 and $1.00, respectively, subject to minimum charges. The current page confirms those prices but does not independently establish the company’s historical comparison.

The expanded workflow complements Together AI’s separate endpoint-level A/B testing for production models. One fine-tuning capability is still on the roadmap: Together AI described deployment of intermediate LoRA adapters while a run is underway as coming soon, with initial support planned for GLM-5.3.

More news