Split or Stay Whole? Disaggregated vs. Monolithic LLM Serving on VMware Cloud Foundation
Broadcom, Friday, July 31st, 2026
Broadcom benchmarks a 120-billion-parameter Nemotron both ways on VCF with NVIDIA Run:ai and Dynamo.
Broadcom ran a 120-billion-parameter Nemotron model on VMware Cloud Foundation using vSphere Kubernetes Service, NVIDIA Run:ai, and NVIDIA Dynamo, serving it both disaggregated and monolithic.
The combination had not been documented anywhere, so the team measured it and wrote up the results. Disaggregated serving splits prefill and decode across separate resource pools, which changes both throughput and the failure profile.
The post reports where each approach wins and notes a finding the team was not looking for. It is aimed at enterprises evaluating on-premises LLM inference on existing VMware estates.