Next upHack for Humanity: San Francisco (powered by Google Gemini)
News

Thinking Machines releases Inkling-Small, a leaner open-weight sibling of its flagship

Thinking Machines Lab released Inkling-Small on July 30, 2026, an open-weight model scoring 40 on Artificial Analysis's Intelligence Index at a third of its flagship's output price.

D
Jul 30, 2026 · 1 min read

Thinking Machines Lab released Inkling-Small on July 30, 2026, an open-weight model a quarter the size of its flagship that nearly matches it on an independent benchmark while costing a third as much to run.

Artificial Analysis, an AI benchmarking firm, scored Inkling-Small 40 on its Intelligence Index versus 41 for the full Inkling, the 975-billion-parameter model the lab shipped July 15. Output pricing is $1.20 per million tokens against $4.05 for the larger model.

Inkling-Small is a multimodal Mixture-of-Experts model, 276 billion total parameters with 12 billion active per query, that takes text, image and audio input and supports a 1-million-token context window. It was trained on Nvidia GB300 NVL72 systems, and the full weights, in BF16 and NVFP4-quantized form, are on Hugging Face under an Apache 2.0 license, with fine-tuning through the lab’s Tinker API.

Thinking Machines reported strong internal scores, including 80.2% on SWE-Bench Verified and 95.5% on AIME 2026. Those benchmarks are the lab’s own and have not been independently verified.

The release fits a broader pattern: shrink a flagship into a cheaper, permissively licensed model that keeps most of its capability. For Thinking Machines, a near-parity small model under Apache 2.0 is a direct pitch to developers weighing open weights against closed commercial APIs. The open question is whether a one-point Intelligence Index gap holds up across the messier real-world tasks benchmarks don’t capture.

More news