The hybrid cooling approach that has become the default in AI infrastructure is not a bridge to the future, it is a trap. Maintaining separate cooling systems for GPUs and storage means paying for two expensive infrastructures while capturing the benefits of neither. As Hardeep Singh of Solidigm notes, you end up with the worst of both worlds: the capital cost of liquid cooling plus the operational drag of air-cooled components that are physically starved of airflow by the very plumbing meant to cool the system.
For infrastructure leaders, this is not an abstract engineering problem. It directly affects GPU utilization. When storage drives throttle under thermal load because bulky cold plates and hoses block the fans that are supposed to cool them, model serving efficiency suffers. Techniques like KV cache offload make this connection explicit: slow storage means a slow model. The race to scale AI is no longer just about who has the most GPUs, but about who can keep them cool, and that now includes keeping storage cool within the same liquid loop. Every component in the rack must operate natively within a shared cooling architecture, or the system will fragment into inefficiency.
The solution requires rethinking storage from the ground up. Traditional SSD designs assume airflow and place components on both sides of a thermally insulated PCB. Neither assumption holds in a fanless, liquid-cooled environment. Solidigm has worked with NVIDIA to address challenges like hot-swap capability and single-side cooling, ensuring that storage does not demand more coolant volume than GPUs can afford to lose. This is not optional engineering; it is the difference between a system that scales and one that throttles. As Singh puts it, if storage is not designed for liquid cooling, it will either underperform or consume coolant that should be going to compute.
The industry is moving toward standards through SNIA and the Open Compute Project, but the practical takeaway for anyone building AI infrastructure today is blunt: do not treat storage as a passive subsystem that can be bolted on after the cooling design is complete. Storage is now an active participant in system-level thermal management, serviceability, and GPU utilization. The design rules have changed. The question is whether your architecture has caught up.
