Cohere ties a three-part data mix to multilingual reasoning
Cohere released gated Tiny Aya L2-Thinker weights and multilingual reasoning data with research on a mix of English reasoning, translated traces and multilingual instructions.
Cohere published research tying a three-part training-data mix to in-language reasoning. Alongside the study, it released gated Tiny Aya L2-Thinker model weights and a public multilingual reasoning dataset. Researchers can inspect those model and data artifacts, but the reported evaluation gains remain Cohere’s results and have not been independently reproduced.
The Tiny Aya L2-Thinker model repository describes a 3.35-billion-parameter model with a combined 32,000-token input-output context length. Cohere Labs licenses it under CC-BY-NC and applies its Acceptable Use Policy. Access to the files, however, requires accepting the repository conditions. The companion dataset contains translated reasoning traces in 44 languages, pairing target-language prompt, thinking and answer fields with the English originals. Its card currently lists 286,376 rows and a 1.24 GB file.
The paper’s supervised fine-tuning recipe draws on three sources. English reasoning traces provide task-solving behavior. English traces translated into 44 non-English languages teach the model to reason in the prompt language. About 4.9 million non-reasoning instruction examples across all 67 Tiny Aya languages reinforce language alignment and transfer, using empty thinking blocks and direct-answer markers instead of reasoning traces.
For the English component, the paper describes a base mix of about 1.7 million examples generated by gpt-oss-120b: roughly 660,000 math, 170,000 science and 890,000 general-reasoning samples. The final model adds about 75,000 Dolci-Think-SFT-32B samples and 380,000 Open-Thoughts-114K samples. Cohere says command-a-translate handled the languages it supported, with DeepSeek-V3 used for the remaining translations.
In an ablation averaged across five benchmarks, Cohere reports that English reasoning data alone delivered 36.7% task accuracy and a 12.8% in-language reasoning rate. Adding multilingual reasoning raised those measures to 45.3% and 86.1%, respectively, according to the company. Adding multilingual non-reasoning data increased them to 57.4% and 96.4%. The authors also report that the final Tiny Aya L2-Thinker exceeded a 93% in-language reasoning rate across 60 languages and six benchmarks. These are author-reported evaluations, not independent measurements.
The evaluation uses FastText to identify the language of reasoning traces, with GlotLID as a fallback. For an open-ended benchmark, the researchers used GPT-4.1 as an automated judge across four rubrics. For context on another measurement approach, Cohere has also published work on retrieval evaluation. The authors say they did not conduct human evaluation of the traces’ quality, fluency or cultural appropriateness. They also note that the reported language rates depend on a prompt instruction used during training and caution that translated supervision can carry English reasoning styles and cultural framing into other languages.
The release records disagree on the dataset total. The current dataset card lists 286,376 rows, while Table 3 of the paper totals 286,388 samples. Neither first-party source explains the 12-row difference.
More news

Cohere introduces RCP-nDCG@10 retrieval metric

Cohere launches Embed 5 Pro and Fast for shared-index retrieval

Cohere releases North Small Translate, an open-weight translation model
