Two weeks after Thinking Machines introduced Inkling, the startup led by former OpenAI CTO Mira Murati is already back with a smaller sibling that deserves serious attention. Inkling-Small is not just a scaled-down curiosity; it is a 276-billion-parameter model that scores within a single point of the 975-billion-parameter flagship on the Artificial Analysis Intelligence Index while using only 12 billion active parameters per token. That is a remarkable efficiency gain, and it arrives with an Apache 2.0 license, full weights on Hugging Face, and a price tag that undercuts the larger model by half during the launch window. For enterprises that have been watching the open-weight space with cautious interest, this is the moment to stop window-shopping. The question is no longer whether open models can compete; it is whether your infrastructure can take advantage of one that finally fits.
What makes Inkling-Small compelling is not the raw benchmark scores, though they are impressive. It is the practical math. The model beats Inkling on SWE-bench Verified, Terminal Bench, and several reasoning tasks, yet it requires roughly a quarter of the total parameters and a fraction of the active compute. That translates into lower inference costs, easier capacity planning, and a deployment footprint that fits on four B300 GPUs or eight H200s. That is still not a laptop, but it is a materially smaller ask than the 3.5-times-larger flagship. For a company with some GPU budget but not a bottomless one, this changes the calculus. You are no longer choosing between a frontier-level model you cannot afford to run and a toy that cannot do the job. You are choosing between two capable open models, one of which costs significantly less to operate. That is a good position to be in, and it is a direct result of Thinking Machines treating model development like a repeatable pipeline rather than a series of one-off miracles. The contrast with Moonshot's custom license on Kimi K3 is telling; Apache 2.0 removes the legal friction that often stalls enterprise adoption before a single token is processed.
But here is the honest tradeoff. Inkling-Small is not uniformly better. Its factual knowledge lags the flagship, with a negative score on AA Omniscience and a weaker showing on τ³-Banking. That means for high-stakes tasks where accuracy is non-negotiable, you still need retrieval-augmented generation, verification layers, and human review. The model is excellent for coding assistants, tool-use systems, document analysis, and multimodal workflows, but it is not a drop-in replacement for a knowledge-intensive application without guardrails. That is not a flaw so much as a design choice, and it is one that Thinking Machines has been transparent about. The variable reasoning effort feature gives developers a dial to trade latency for quality, which is a level of control that most proprietary APIs simply do not offer. If you are building a product where you can tolerate some hallucination risk in exchange for lower cost, this model is a strong candidate. If you are building a medical diagnosis tool, you still need to do your homework.
The broader signal here is the one worth watching. Thinking Machines has demonstrated that it can compress a flagship model into a quarter of the size without collapsing performance, and it did so in two weeks. That suggests a mature training pipeline, one that can iterate quickly and ship improvements on a cadence that legacy vendors cannot match. The company is not just releasing models; it is building a system for producing them. For enterprises, that means the choice is not between Inkling-Small and Inkling today. It is between adopting a model from a vendor that is clearly learning how to make open-weight AI more accessible, or waiting for the next release that will inevitably be smaller, cheaper, and possibly better. The specific detail to watch is how quickly third-party quantization and fine-tuning communities adopt this model. If the open-source ecosystem can shrink it further without a meaningful drop in quality, the window for proprietary models narrows considerably. That is the real story here, and it is just getting started.
