generative AI for data analysis

Smarter attention, faster inference: IndexCache cuts redundant work in AI models

Introducing **IndexCache**, a groundbreaking sparse attention optimizer designed to enhance the efficiency of long-context AI models.

3 min readVentureBeat
Smarter attention, faster inference: IndexCache cuts redundant work in AI models

The real story here isn't the clever trick IndexCache uses, it's that the AI industry has been paying a tax it didn't need to pay. Researchers at Tsinghua University and Z.ai found that DeepSeek Sparse Attention models repeatedly compute the same thing across adjacent layers, with 70 to 100 percent of selected tokens staying identical. That is wasted work, and IndexCache eliminates it by letting most layers simply reuse what the layer before already figured out. For anyone running long-context models at scale, this is the kind of optimization that turns a server cost problem into a manageable line item.

What this means for your infrastructure is straightforward. If you are serving a 30-billion-parameter model at 200,000 tokens of context, IndexCache cuts prefill latency from 19.5 seconds to 10.7 seconds. That is nearly halving the time your users wait for a first response. Decode throughput jumps from 58 tokens per second to 86. These are not theoretical gains from a toy benchmark; they come from production-scale tests on the GLM-4.7 Flash model and preliminary results on the 744-billion-parameter GLM-5. The technique works without retraining, so you can apply it to models you already have deployed. The open-source patches for vLLM and SGLang mean your engineering team can integrate it with minimal configuration changes.

The more interesting implication is what this says about where the field is headed. IndexCache solves a problem that only exists because the original architecture was designed without thinking hard enough about inference cost. The DSA indexer is lightweight compared to full attention, but it still scales quadratically with sequence length. That is a bottleneck the researchers found because they looked at the whole pipeline, not just the headline metric of model quality. Their solution, sharing indices across layers, is elegant precisely because it exploits a property of how these models actually behave, not how we assume they behave. Co-author Yushi Bai put it plainly: future foundation models will likely be architected with downstream inference constraints in mind from the start. That is not a prediction about some distant future; it is a design principle that IndexCache already proves is viable today.

For teams running RAG pipelines, document analysis workloads, or agentic systems that chew through long contexts, the practical takeaway is clear. You can deploy IndexCache now, cut at least 20 percent from your deployment costs, and match or even exceed your baseline quality scores. The AIME 2025 math reasoning benchmark actually improved by 1.6 points after removing 75 percent of the indexers. That is the kind of result that makes you question what other hidden inefficiencies are waiting to be found. The industry has been throwing compute at a problem that was partly architectural. IndexCache shows that sometimes the smartest optimization is simply not doing the work you never needed to do.

From VentureBeat

Processing 200,000 tokens through a large language model is expensive and slow: the longer the context, the faster the costs spiral. Researchers at Tsinghua University and Z.ai have built a technique called IndexCache that cuts up to 75% of the redundant computation in sparse attention models, delivering up to 1.82x faster time-to-first-token and 1.48x faster generation throughput at that context length.

Read the original at VentureBeat