On September 30, 2026, CoreWeave announced that NVIDIA Vera Rubin NVL72 was available on CoreWeave Cloud and named Cognition as its first customer running production workloads on the system. Cognition reported separate throughput results for SWE-2 inference and reinforcement learning.

Cognition’s Vera Rubin workloads and results

CoreWeave said the Cognition cluster was stood up in early September. The company describes Cognition as running production workloads on Vera Rubin NVL72; the specific SWE-2 result concerns inference, or running a model to generate responses, rather than a reported training run.

Cognition’s engineers reported up to 4.8× total token throughput for SWE-2 inference on Vera Rubin NVL72 compared with an NVIDIA GB200 NVL72 baseline. For a separate reinforcement-learning workload, Cognition reported 3.8× output-token throughput. These figures describe different workloads and metrics, so they are not interchangeable.

What the two throughput figures measure

The 4.8× result is for total token throughput in SWE-2 inference, against a specified GB200 NVL72 baseline. The 3.8× result is for output-token throughput in reinforcement learning. Keeping those workloads and measures distinct is key to understanding the reported results: neither figure is a general performance claim for every AI task.

The rack behind Cognition’s workloads

A Vera Rubin NVL72 rack appears in a data-center setting as equipment is unloaded and technicians handle the infrastructure.

CoreWeave’s Vera Rubin NVL72 rack configuration pairs 72 Rubin GPUs with 36 Vera CPUs. On September 16, 2026, CoreWeave announced multi-rack Vera Rubin NVL72 clusters on its cloud. It named Cognition’s production workloads in a separate announcement on September 30. NeoTeo covered the earlier cluster announcement in an earlier report.