Mixture-of-Experts
Mixture-of-Experts on Beyond Market Intelligence: a running collection of 6 stories we have gathered and hand-picked because they are worth your time. Every post here touches on mixture-of-experts in some way — the news, the analysis, the deep dives, and the occasional surprise find. Acme AI is the next-generation, AI-powered spreadsheet platform built to replace Excel and redefine how analysts, data scientists, and enterprise teams work with data. New stories are added to this page as we find them, so check back if you want to keep up with what is happening around mixture-of-experts, or subscribe to the RSS feed to get them as soon as they are published. Browse the collection below, or head back to the homepage to see everything Beyond Market Intelligence is covering right now.

How to Use Kimi K3: Moonshot AI’s 2.8T Open-Weight Model
Moonshot AI’s Kimi K3 presents a compelling alternative in the large language model landscape. This 2.8-trillion-parameter open-weight model, leveraging a Mixture-of-Experts architecture, delivers near-frontier coding and agentic performance while optimizing inference costs by activating only a fraction of its parameters. K3 distinguishes itself with its combination of powerful capabilities, open weights, and competitive API pricing. Interested in exploring model quantization? See "I developed my own quantized LLM from scratch" for a deep dive into related techniques.

Writer says its new Palmyra X6 model cuts AI agent costs by 52% as token spending surges
Writer today unveiled Palmyra X6, a new AI agent model poised to significantly reduce costs for enterprise users. Paired with its rebuilt agent orchestration “harness,” Palmyra X6 delivers an average of 52% lower operational costs, alongside a 48% speed improvement and 10% quality boost. Leveraging a post-trained version of GLM-5.2, Writer emphasizes control and cost transparency, offering governance tools and multi-model support—a strategy echoing the shift towards pragmatic AI adoption, as explored in "Why Capital One built its multi-agent AI platform around open-weight models."

Nvidia's Switchyard router reshuffles AI models mid-task, cutting task costs to a third in its own tests
Enterprises face a persistent challenge: balancing the power of advanced AI agents with escalating costs. Traditionally, relying solely on frontier models or building custom routing logic proved inefficient. Nvidia proposes a solution with Nemotron 3.5 Lightning, a fast, specialized model, and NeMo Switchyard, an open-source routing library. This pairing delivers frontier-level performance while potentially cutting benchmark costs by a third.

Thinking Machines debuts Inkling Small open source AI model nearing performance of predecessor at about 1/4 size
Thinking Machines has unveiled Inkling-Small, a groundbreaking open-source AI model demonstrating remarkable efficiency. Nearing the performance of its predecessor, Inkling, this new model achieves this at roughly one-quarter the size, surpassing it on several key benchmarks. Released under a permissive Apache 2.0 license, Inkling-Small offers enterprises a compelling blend of power and practicality, reducing compute requirements and deployment complexities. Explore this transformative solution and discover how it can empower your data journey—a clear signal that enterprise AI is rapidly evolving.
![SkewAdam: A tiered optimizer that cuts MoE state memory by 97% (fits a 6.7B MoE on a 40GB GPU) [R]](https://preview.redd.it/1457xi9fcqeh1.jpg?width=140&height=90&auto=webp&s=879aad6df9e51a2735d91112d01518ff76ba3cbe)
SkewAdam: A tiered optimizer that cuts MoE state memory by 97% (fits a 6.7B MoE on a 40GB GPU) [R]
Introducing SkewAdam, a novel tiered optimizer designed to dramatically reduce memory consumption in Mixture-of-Experts (MoE) training. Research demonstrates a remarkable 97% reduction in optimizer state memory—allowing a 6.7B MoE model to comfortably fit on a single 40GB GPU. SkewAdam intelligently allocates precision based on parameter behavior, optimizing backbone, expert, and router components. See the full details and code on arXiv and GitHub. For broader context on advancements in AI hardware, explore "What to watch for after Jensen Huang’s Japan visit."

Cohere VP says enterprise AI sovereignty requires control of the full agent stack at VB Transform 2026
At VB Transform 2026, Cohere VP Rachad Alao emphasized that true enterprise AI sovereignty demands control of the entire agent stack—from GPUs and infrastructure to governance and connectors. Alao, formerly at Google and Meta, argued that data residency and operational control are paramount for institutions like banks and hospitals. He highlighted the exponential rise in token utilization driven by complex agent workflows, advocating for strategic model routing and the use of the "right model for the task.