Stop Recomputing Context: Introducing ContextCube

← Back to blog

Every time an inference server computes a long prompt, it produces something valuable: the KV cache. Too often, that work is thrown away minutes later, then computed again on the next server that sees the same context.

Today at Tech Week Singapore 2026, we launched ContextCube to change that. ContextCube is a purpose-built KV cache appliance that gives your entire inference cluster one shared, persistent pool of context. GPUs fetch what has already been computed, instead of recomputing it.

Introducing ContextCube: A Shared KV Cache Pool for Your Whole Inference Cluster

The problem: good context, stuck in the wrong place

The KV cache holds the attention keys and values a model computes while reading a prompt. Reusing it is far cheaper than recomputing it, and that matters more as prompts get longer.

In most clusters today, though, the KV cache lives locally: in one GPU's HBM, or at best in that server's DRAM or NVMe. That creates two familiar problems:

  • Eviction. Long prompts fill local memory fast, so cached context is pushed out before it can be reused.
  • Stranding. When a user's next request is routed to a different server, the cache sitting on the first server is out of reach.

The result is the same context computed again and again. Time to first token goes up, and GPU cycles that should generate new tokens are spent on prefill instead.

How ContextCube works: a shared Layer 3.5

ContextCube is designed on the NVIDIA CMX inference-storage reference architecture. It adds a new, shared tier to the memory hierarchy, which we call Layer 3.5, alongside GPU HBM, DRAM and local NVMe.

The division of labour is simple. Active computation stays on the GPUs. Reusable context moves into a large, cluster-wide cache that every participating node can reach.

When a request arrives, the inference node checks ContextCube for matching KV cache. If it finds a match, it retrieves the cached KV over RDMA instead of recomputing it. Because the cache is shared, it does not matter which server handled the earlier request.

Each node keeps its own HBM, DRAM and NVMe; ContextCube adds one cache tier they all share.

What's inside

Each ContextCube appliance is built for fast, DPU-driven data movement, and you can start with just one.

Because ContextCube works with the inference engines and KV cache software teams already use, it slots into an existing stack rather than replacing it.

Built for context-heavy workloads

ContextCube pays off wherever the same context comes back again and again:

  • AI agents that carry long tool histories and system prompts across many steps
  • Coding assistants that repeatedly read the same repositories and files
  • Multi-turn conversations, where each new turn builds on everything before it
  • Long-document processing, where many questions are asked of one large document
  • Retrieval-augmented generation, where popular source passages are retrieved over and over
  • Large-scale "token factories" serving high volumes of requests across many GPUs

For all of these, less recomputation means a faster time to first token, more GPU cycles for generation, and higher throughput and concurrency from the same hardware.

Why we built it

"As context windows grow and AI agents move into production, the same context is being computed again and again on every server. ContextCube gives the whole cluster one place to keep that work, so GPUs can spend their time generating new tokens instead of rebuilding old ones." Jenvik Li, Founder and CEO, TuringData

Get started

ContextCube is available to order now. To discuss your cluster, workloads and deployment, contact our team.

*Actual performance depends on configuration and workload. NVIDIA and BlueField are trademarks and/or registered trademarks of NVIDIA Corporation in the United States and other countries.

Find out how much GPU utilization you're leaving on the table. Schedule a free 30-minute architecture review with our experts.