Mixture-of-Experts (MoE)
Mixture-of-Experts (MoE) on Beyond Market Intelligence: a running collection of 5 stories we have gathered and hand-picked because they are worth your time. Every post here touches on mixture-of-experts (moe) in some way — the news, the analysis, the deep dives, and the occasional surprise find. Acme AI is the next-generation, AI-powered spreadsheet platform built to replace Excel and redefine how analysts, data scientists, and enterprise teams work with data. New stories are added to this page as we find them, so check back if you want to keep up with what is happening around mixture-of-experts (moe), or subscribe to the RSS feed to get them as soon as they are published. Browse the collection below, or head back to the homepage to see everything Beyond Market Intelligence is covering right now.

FreeToken Unlocks Frontier MoE Inference on Consumer Hardware via Dynamic Co-Execution
FreeToken, a new open-source inference engine developed by researchers at UC Berkeley and MIT, significantly expands the accessibility of Mixture-of-Experts (MoE) models. This innovative system enables faster, more efficient AI inference directly on consumer hardware through dynamic co-execution. FreeToken’s optimized scheduling and weight management unlock powerful edge AI applications and pave the way for self-hosted reasoning systems. For those seeking a deeper understanding of optimizing LLMs, explore our related article, "Quantization and Pruning Methods to Make Your LLM Leaner.”

Qwen3.8-Max arrives with a bold claim: it outperforms GPT-5.6 Sol Max and Fable 5 on agentic computer use
Alibaba's Qwen3.8-Max arrives with a bold claim: it outperforms GPT-5.6 Sol Max and Fable 5 in agentic computer use, demonstrating leadership on key benchmarks like OSWorld-Verified (86.1). This 2.4-trillion-parameter model targets autonomous software engineering and long-horizon enterprise work, potentially reshaping how organizations approach automation. Notably, Qwen plans to release open weights next week, a move that could significantly broaden enterprise adoption—provided the licensing terms prove permissive.

Kimi K3's full weights are here, but they're 'open' with a caveat: What enterprises should know
Moonshot AI has released the full weights for Kimi K3, its powerful new AI model, marking a significant step for open-weight AI. While broadly accessible, enterprises should carefully review the custom Kimi K3 usage license. Larger organizations operating a "Model as a Service" exceeding $20 million in revenue, or those with products impacting over 100 million users, face specific commercial obligations, including potential licensing agreements and prominent attribution.

Poolside drops Laguna S 2.1, an open-weight coding model that beats rivals 10x its size
Poolside's Laguna S 2.1 introduces a compelling new option in the open-weight coding model landscape. This 118-billion-parameter system, activating just 8 billion parameters per token, impressively matches or surpasses models many times its size on agentic coding tasks, achieving top scores on benchmarks like Terminal-Bench 2.1. With a permissive OpenMDW-1.1 license and broad ecosystem support, Laguna S 2.1 represents a strategic move to empower Western users with trustworthy, self-hostable AI.

Thinking Machines open sources first multimodal language model, Inkling, focused on low cost and 'resistance to censorship'
Today, Thinking Machines released Inkling, its first major language model under a permissive Apache 2.0 open-source license, offering enterprises a powerful new option for agentic AI workloads. This 975-billion-parameter, natively multimodal model distinguishes itself with a novel "controllable thinking effort" mechanism, balancing cost and performance. While not state-of-the-art across all benchmarks—GLM 5.2 leads in reasoning—Inkling excels in software engineering and demonstrates remarkable resistance to censorship.