CoreWeave Brings NVIDIA Vera Rubin NVL72 and Vera CPUs to AI Cloud
CoreWeave is expanding its AI-optimized cloud with next-generation NVIDIA compute, networking, and software. The companies’ nearly decade-long engineering collaboration continues to support strong returns across multiple generations of deployments, with CoreWeave now bringing NVIDIA Vera Rubin infrastructure into production.
At CoreWeave Fully Connected in San Francisco, CoreWeave will showcase the NVIDIA Vera Rubin NVL72 system with Spectrum-X 102.4T Ethernet networking. Devin AI Cognition, an applied AI lab that powers software engineers, is the first customer to run production workloads on Vera Rubin.
CoreWeave is also making NVIDIA Vera, the first CPU built for AI agents, available on CoreWeave Cloud. In addition, CoreWeave has launched CoreWeave Forge, a connected environment for training, evaluating, and improving models and agents on NVIDIA-accelerated computing.
“NVIDIA Accelerated Computing delivers generational value,” said Ian Buck, vice president of hyperscale and high-performance computing at NVIDIA. “Even now, as CoreWeave brings the Vera Rubin NVL72 into production, nearly a decade after the launch of Volta, CoreWeave’s NVIDIA V100 GPUs are running customer workloads. This is the strength of the NVIDIA platform: an infrastructure that continues to generate revenue for years and the flexibility to put the right GPU in the right workload.”
Vera Rubin NVL72 delivers up to 4.8x higher token throughput for AI coding workloads
Cognition runs Devin training, reinforcement learning, and production inference on CoreWeave. In nine months, the company scaled to thousands of GPUs on CoreWeave to support its Cognition inference workloads.
Earlier this month, CoreWeave received its first Vera Rubin NVL72 production rack. Cognition then benchmarked Vera Rubin inference performance against the GB200 NVL72 baseline using real-world software engineering workloads. The benchmark used a subset of tasks sampled from FrontierCode and an AI agent deployed to solve them.
In early testing, Cognition confirmed that Vera Rubin NVL72 improves aggregate token throughput by up to 4.8x for SWE-2 inference workloads compared with GB200 NVL72. For Devin, this means faster real-time code generation and more responsive multi-step inference.
“Agent coding is a complex workload: long contexts, high concurrency, high token volumes, and cost per token that determines what you can ship,” said Silas Alberti, a member of Cognition’s founding team. “Having everything on one platform and having NVIDIA and CoreWeave engineers working with us on tough problems is more important than a single specification.”
NVIDIA Vera Rubin NVL72 now available on CoreWeave Cloud
CoreWeave has announced the availability of NVIDIA Vera Rubin NVL72 on CoreWeave Cloud, making it one of the first cloud providers to put the platform in customers’ hands.
Early Access customers can use the performance of NVIDIA’s full-stack AI Factory platform on CoreWeave Cloud. CoreWeave launched a production Vera Rubin cluster for Cognition within days, enabled by collaboration between NVIDIA and CoreWeave across the stack, from infrastructure to the tokens provided.
Capacity can be managed through CoreWeave Kubernetes Service, SUNK, CoreWeave Mission Control, CoreWeave Sandboxes, and CoreWeave Inference.
NVIDIA Vera CPUs accelerate AI agent sandbox environments
Agentic AI places pressure on infrastructure from two directions. Serving agents requires large-scale, low-latency computing, while improving agents after training requires thousands of isolated environments to run simultaneously.
NVIDIA Vera CPUs are purpose-built for agent workloads. A key performance measure for agentic AI is how many separate environments can run concurrently and how consistently each environment maintains performance as demand increases.
CoreWeave’s Vera deployment includes 128 CPUs and 11,264 cores in a single rack. With one core sufficient for each environment, the rack can support more than 11,000 concurrent environments. CoreWeave Sandbox isolates these environments in hardware and runs them in parallel with training jobs supported by Spectrum-X Ethernet switches and BlueField-4 DPUs. This enables secure, high-performance agent communication with low latency.
Testing showed that CoreWeave achieved more than 3x faster agent sandbox startup times on NVIDIA Vera CPUs. The improvement accelerates sandbox use, reinforcement learning execution, agent tooling, and model evaluation, allowing AI teams to run code in isolated environments on CoreWeave. On the terminal benchmark, CoreWeave also saw a 1.7x performance improvement on Vera CPUs across all passing tasks.
CoreWeave Forge connects production, evaluation, and AI training
Models and agents improve through continuous feedback loops. Operational behavior informs the next training run, while each evaluation strengthens the next version. Traditionally, these workflows have been split across tools from different vendors, creating signal loss during handoffs.
CoreWeave Forge combines Weights & Biases, OpenPipe’s post-training expertise, and the open-source marimo notebook project in one connected environment for continuous model and agent improvement. Forge remains open across models, frameworks, and clouds.
New features and enhancements include:
- CoreWeave ARIA — Now generally available, ARIA helps users learn, explore, code, and iterate across the AI loop. It analyzes runs, suggests experiments, recommends code changes, saves them to GitHub, and provides actionable evidence to help identify why results changed and determine the next experiment.
- CoreWeave Agent Lens — This new service turns operational agent observability into insights for continuous improvement. It improves fault detection by 20% and solves problems at half the cost by turning tens of millions of operational agent traces into remediation insights.
- CoreWeave Sandbox — Now generally available, Sandbox lets users run agents, tool invocations, reinforcement learning, and evaluations in isolated CPU or GPU environments. Workloads can run on serverless infrastructure or on infrastructure already used for training, providing a fresh environment for each agent tool invocation, RL execution, or evaluation.
- Post-training for improved model quality, latency, and cost — Teams can use production signals without a dedicated training cluster. Fine-tuning with serverless monitoring and serverless RL let users experiment with their own training recipes. Serverless RL trains 1.4x faster at 40% lower cost than self-managed setups.
NVIDIA Dynamo powers CoreWeave’s Managed Inference Service. The open-source inference framework for AI factories also supports RL rollout, which is currently in private preview. RL rollout loads new checkpoints into a live deployment during execution, allowing reinforcement learning to continue without redeployment while maintaining the same post-training inference efficiency as production.
Canva, Capital One, and MasterClass were among the first companies to build on Forge.
NVIDIA Nemotron, the open model, gives teams on Forge a direct path to customize and deploy inference and multimodal models for agent workflows.
AI infrastructure for startups, AI labs, and global enterprises
AI labs, AI-native companies, and global enterprises use the co-designed NVIDIA and CoreWeave platform to move from prototype to production faster.
In healthcare, Ennoble Care, a home health provider serving approximately 50,000 advanced Medicare patients across 15 states, selected CoreWeave to run clinical AI inference. The company can scale AI agents for clinical documentation, decision support, and back-office automation with reserved NVIDIA RTX PRO 6000 GPU capacity on CoreWeave Kubernetes Service.
CoreWeave has also achieved leading MLPerf results in training and inference. It is the only cloud provider to hold platinum rankings in SemiAnalysis ClusterMAX 1.0, 2.0, and 3.0, serving 9 of the 10 leading AI labs.
Together, NVIDIA and CoreWeave are providing customers with a platform to turn experimental agents into production systems that create software, support clinicians, and perform useful work in the real world.
Join and learn more about NVIDIA Sessions, Demos, and Workshops at CoreWeave Fully Connected.
Source: blogs.nvidia.com


