
VentureBeat
Stop adding more GPUs: Weka's new storage platform reduces load by caching 100% of an AI model's pre-calculated tokens
GPU memory is rapidly becoming the primary bottleneck in production AI, particularly as models demand longer context windows. Weka’s new storage platform directly addresses this challenge, offering a transformative approach that extends GPU capacity with cost-effective flash storage. Through its NeuralMesh 6 software and Wekapod 3 hardware, Weka’s Augmented Memory Grid caches 100% of pre-calculated tokens, eliminating redundant computations and significantly reducing inference costs.














![Tri-Net v2: Open-source implementation of our Scientific Reports paper on unified skin lesion and symptom-based monkeypox detection [R]](https://preview.redd.it/vwax5ludzheh1.png?width=140&height=79&auto=webp&s=25929233532a0110f28de21f8e7a57634c6f791b)

























