Thinking Machines releases Inkling-Small, a leaner open-weight sibling of its flagship
Thinking Machines Lab released Inkling-Small on July 30, 2026, an open-weight model scoring 40 on Artificial Analysis's Intelligence Index at a third of its flagship's output price.
Thinking Machines Lab released Inkling-Small on July 30, 2026, an open-weight model a quarter the size of its flagship that nearly matches it on an independent benchmark while costing a third as much to run.
Artificial Analysis, an AI benchmarking firm, scored Inkling-Small 40 on its Intelligence Index versus 41 for the full Inkling, the 975-billion-parameter model the lab shipped July 15. Output pricing is $1.20 per million tokens against $4.05 for the larger model.
Inkling-Small is a multimodal Mixture-of-Experts model, 276 billion total parameters with 12 billion active per query, that takes text, image and audio input and supports a 1-million-token context window. It was trained on Nvidia GB300 NVL72 systems, and the full weights, in BF16 and NVFP4-quantized form, are on Hugging Face under an Apache 2.0 license, with fine-tuning through the lab’s Tinker API.
Thinking Machines reported strong internal scores, including 80.2% on SWE-Bench Verified and 95.5% on AIME 2026. Those benchmarks are the lab’s own and have not been independently verified.
The release fits a broader pattern: shrink a flagship into a cheaper, permissively licensed model that keeps most of its capability. For Thinking Machines, a near-parity small model under Apache 2.0 is a direct pitch to developers weighing open weights against closed commercial APIs. The open question is whether a one-point Intelligence Index gap holds up across the messier real-world tasks benchmarks don’t capture.
More news

AWS releases six open-source Hugging Face deployment skills for SageMaker

Google Research releases MilleMiglia logistics benchmark generator

AWS launches AgentCore Runtime V2 with elastic memory and snapshot starts
