CoreWeave Delivers NVIDIA Vera Rubin NVL72 Performance at Production Scale, Starting With Cognition
HPCwire, Wednesday, September 30th, 2026
CoreWeave puts NVIDIA Vera Rubin NVL72 into production, with Cognition reporting up to 4.8x higher token throughput per GPU.
CoreWeave announced availability of NVIDIA Vera Rubin NVL72 on its cloud, with Cognition the first customer anywhere running production workloads on the system.
Cognition measured up to 4.8x higher total token throughput per GPU for SWE-2 inference at matched interactivity and 3.8x higher output throughput for reinforcement-learning workloads, meaning more concurrent Devin sessions per GPU at lower cost.
CoreWeave stood up the cluster in early September and had workloads running within days of handover, with each rack pairing 72 Rubin GPUs and 36 Vera CPUs with 20.7 TB of HBM4 and sixth-generation NVLink.