Together AI launches Tev1-4B experimental classifier
Together AI launched hosted classifier Tev1-4B-experimental and outlined how to fine-tune Qwen3.5-4B for structured decision tasks.
Together AI has launched Tev1-4B-experimental on its serverless platform and published a workflow for training a related classifier on Qwen3.5-4B. The small model takes a supplied state, a question and between two and 24 labeled options, then returns a single option letter that an application can map to a semantic key.
Together positions that constrained output for structured decisions including customer-support routing, yes-or-no comprehension, policy checks, research-paper categorization and sentiment classification. These are potential uses, not independently measured production results. The model card says Tev1 is inspired by Jev but is not a non-autoregressive Jev runtime; it keeps Qwen’s standard next-token language-model head.
The published recipe starts by cloning the Tev1 repository, installing locked dependencies with uv, downloading source datasets with fetch_sources.py, and sampling and normalizing them with build_all.py. The retained recipe records 37,840 unique training examples and 4,568 validation examples across natural-language inference, reading comprehension, banking intent, news classification, sentiment, programmatic policies, routing and research taxonomy data.
One figure in Together’s post does not match that record. The dataset table totals 37,840 training examples, in line with the repository, but a later cost sentence refers to 38,340 examples.
Without --launch, the training command runs as a preview. Adding the flag uploads train.jsonl and dev.jsonl to Together, starts a billed fine-tuning job, and returns a job ID that can be tracked through the command-line interface or dashboard. The repository describes standard LoRA supervised fine-tuning and saves a proposed configuration with rank 8, one epoch, a 5e-5 learning rate and a 2,048-token limit. It also says the historical run’s settings and uploaded files could not be verified.
Together says the example run cost about $17 and took roughly 25 minutes. Those are company claims: the sources include neither a billing receipt nor a retrievable job record, and the run was not independently reproduced.
After training, the workflow retrieves the output model name and creates a dedicated endpoint on one Nvidia H100 80GB SXM GPU. It stores the endpoint’s model name in the environment configuration and sends decisions through the supplied script. This dedicated deployment path is separate from Together’s description of the prebuilt Tev1 model as serverless. For related deployment context, Together has also added endpoint-level A/B testing for production models.
For inference, Together recommends temperature zero, no more than eight output tokens, disabled thinking, and a system instruction that treats the state as data and allows exactly one option letter.
Together reports 880 correct answers out of 1,000 on its main decision set, 300 out of 300 on a policy-transfer set, and valid single-letter output in all 1,300 evaluations. The company describes these as reused development results rather than independent or untouched benchmarks, and it provides no untuned Qwen3.5-4B baseline.
The model card warns that Tev1 can be wrong and should not be the sole authority for high-impact decisions. It says prompt injection, multilingual behavior, calibration and broad out-of-distribution robustness have not been comprehensively evaluated. The card also says the release license for the fine-tuned weights is still being finalized and that the training mixture has no single blanket dataset license.
More news

xAI launches Team Bots for shared workflows

AWS adds xAI’s Grok 4.7 to Amazon Bedrock

OpenAI adds $5 million and up to $5 million in credits to Lenfest AI program
