generative AI for data analysis

Inkling Small delivers near-flagship performance at a fraction of the size.

Just two weeks after unveiling Inkling, Thinking Machines is back with Inkling-Small, and it's a compelling trade.

4 min readVentureBeat
Inkling Small delivers near-flagship performance at a fraction of the size.

Two weeks after Thinking Machines introduced Inkling, the startup led by former OpenAI CTO Mira Murati is already back with a smaller sibling that deserves serious attention. Inkling-Small is not just a scaled-down curiosity; it is a 276-billion-parameter model that scores within a single point of the 975-billion-parameter flagship on the Artificial Analysis Intelligence Index while using only 12 billion active parameters per token. That is a remarkable efficiency gain, and it arrives with an Apache 2.0 license, full weights on Hugging Face, and a price tag that undercuts the larger model by half during the launch window. For enterprises that have been watching the open-weight space with cautious interest, this is the moment to stop window-shopping. The question is no longer whether open models can compete; it is whether your infrastructure can take advantage of one that finally fits.

What makes Inkling-Small compelling is not the raw benchmark scores, though they are impressive. It is the practical math. The model beats Inkling on SWE-bench Verified, Terminal Bench, and several reasoning tasks, yet it requires roughly a quarter of the total parameters and a fraction of the active compute. That translates into lower inference costs, easier capacity planning, and a deployment footprint that fits on four B300 GPUs or eight H200s. That is still not a laptop, but it is a materially smaller ask than the 3.5-times-larger flagship. For a company with some GPU budget but not a bottomless one, this changes the calculus. You are no longer choosing between a frontier-level model you cannot afford to run and a toy that cannot do the job. You are choosing between two capable open models, one of which costs significantly less to operate. That is a good position to be in, and it is a direct result of Thinking Machines treating model development like a repeatable pipeline rather than a series of one-off miracles. The contrast with Moonshot's custom license on Kimi K3 is telling; Apache 2.0 removes the legal friction that often stalls enterprise adoption before a single token is processed.

But here is the honest tradeoff. Inkling-Small is not uniformly better. Its factual knowledge lags the flagship, with a negative score on AA Omniscience and a weaker showing on τ³-Banking. That means for high-stakes tasks where accuracy is non-negotiable, you still need retrieval-augmented generation, verification layers, and human review. The model is excellent for coding assistants, tool-use systems, document analysis, and multimodal workflows, but it is not a drop-in replacement for a knowledge-intensive application without guardrails. That is not a flaw so much as a design choice, and it is one that Thinking Machines has been transparent about. The variable reasoning effort feature gives developers a dial to trade latency for quality, which is a level of control that most proprietary APIs simply do not offer. If you are building a product where you can tolerate some hallucination risk in exchange for lower cost, this model is a strong candidate. If you are building a medical diagnosis tool, you still need to do your homework.

The broader signal here is the one worth watching. Thinking Machines has demonstrated that it can compress a flagship model into a quarter of the size without collapsing performance, and it did so in two weeks. That suggests a mature training pipeline, one that can iterate quickly and ship improvements on a cadence that legacy vendors cannot match. The company is not just releasing models; it is building a system for producing them. For enterprises, that means the choice is not between Inkling-Small and Inkling today. It is between adopting a model from a vendor that is clearly learning how to make open-weight AI more accessible, or waiting for the next release that will inevitably be smaller, cheaper, and possibly better. The specific detail to watch is how quickly third-party quantization and fine-tuning communities adopt this model. If the open-source ecosystem can shrink it further without a meaningful drop in quality, the window for proprietary models narrows considerably. That is the real story here, and it is just getting started.

From VentureBeat

Just two weeks after Thinking Machines released Inkling, its first open source AI language model, the well-funded startup led by former OpenAI chief technology officer Mira Murati today introduced Inkling-Small without sacrificing much of any performance — and in fact, the new model surpasses its larger predecessor on several benchmarks.

Read the original at VentureBeat