Back Issues/Search Home → Calendar → Archive → Current Issue → Popular →

All issuesVolume 341, Issue 3IT Vendor NewsDell

KV Cache Offload to Object Storage: GPU-Direct and Upstream

Dell, Wednesday, August 19th, 2026

Dell and NVIDIA upstreamed an accelerated NIXL OBJ engine enabling RDMA KV cache offload from vLLM and LMCache to ObjectScale.

Dell Technologies, working with NVIDIA, contributed an accelerated engine for the NIXL OBJ plugin that has been merged upstream, letting vLLM, LMCache and NIXL offload KV cache to Dell ObjectScale over RDMA and land it directly in GPU memory via cuObject. No fork or custom client is required.

Dell measured 837 ms to first token at a 235,000-token context against 11,223 ms when prefill is recomputed, a 13.4 times improvement.

The same offload was 1.3 to 1.5 times faster than running over standard S3. Dell argues the performance reason to keep object storage out of the inference path is now gone.

more →  ·  More from Dell →