NVIDIA Vera CPU Rack 256 Vera CPUs 88 Olympus cores 48U · 100% liquid-cooled Agentic AI & RL
NVIDIA Vera CPU Rack — a dense, liquid-cooled 48U MGX rack with 256 Vera CPUs for agentic AI and reinforcement-learning sandbox environments

The NVIDIA Vera CPU Rack is a dense, fully liquid-cooled 48U system that packs 256 Vera CPUs — built on the MGX modular platform — to run the CPU-side work behind agentic AI: code execution, tool calls and reinforcement-learning sandbox environments at factory scale. This guide explains what Vera is, the specs of the chip and the rack, how it pairs with GPU systems, where it fits in your agentic AI roadmap, and what to plan before you deploy in Europe.

Short answer

The Vera CPU Rack is a CPU-first rack purpose-built for agentic AI and reinforcement learning, where models generate code and queries and CPUs execute the actions, evaluate results and return data across thousands of parallel environments. A single 48U rack integrates 256 Vera CPUs — 22,528 Olympus cores, up to 400 TB of LPDDR5X memory and 64 BlueField-4 DPUs — to sustain more than 22,500 concurrent CPU environments, and it is designed to run alongside GPU rack systems rather than replace them.

Why this matters for infrastructure buyers

Vera CPU Rack is not just another CPU system. It matters if your GPU infrastructure is starting to wait behind CPU-side execution: code sandboxes, tool calls, reinforcement-learning environments, retrieval and agent workflows. If you are planning an AI factory or agentic AI platform in Europe, the key question is not only whether Vera is powerful enough, but whether your facility, cooling, networking and GPU roadmap are ready for this type of rack-scale CPU layer.

What it is

A liquid-cooled CPU rack for the control-heavy, latency-sensitive work of agents and RL — not a GPU trainer, but the execution layer that keeps GPUs fed.

Who should read this

Teams building agentic AI platforms, RL training pipelines and AI factories who need to scale CPU sandbox environments alongside their GPU compute.

Planning agentic AI infrastructure?

SERVER SIMPLY helps you scope and source the CPU and GPU infrastructure behind agentic AI and RL workloads in Europe — with configuration, EU delivery and binding lead times on quote.

Get a configuration & lead time

Why agentic AI needs a new kind of CPU

Agentic AI changed what the CPU does in an AI system. When a model writes code, calls a tool or issues a query, something has to actually run that action, evaluate the result and return data to the model — and that something is a CPU. In reinforcement learning and agent training, this loop repeats across thousands to millions of parallel environments, each needing predictable, low-latency compute. Traditional data-center CPUs, optimized for cores-per-dollar on general workloads, were not designed for that pattern.

That is the gap the Vera CPU Rack targets: a CPU and a rack built specifically for the execution side of agents and RL, so the CPU layer stops being the bottleneck that starves expensive GPU capacity. It is designed to work alongside GPU rack systems such as the Vera Rubin NVL72, not to replace them.

This is also where SERVER SIMPLY fits. We help European AI teams translate this NVIDIA roadmap into practical infrastructure decisions: what can be deployed now, what should be planned for 2026, what cooling and rack requirements are needed, and how CPU and GPU capacity should be balanced for your agentic AI platform.

What is the NVIDIA Vera CPU Rack?

The Vera CPU Rack is a dense, 100% liquid-cooled 48U system built on the NVIDIA MGX modular platform. It integrates 256 Vera CPUs together with 64 BlueField-4 DPUs and Spectrum-X Ethernet for networking and infrastructure services, delivered as a ready-to-deploy rack rather than a set of loose servers. The goal is to remove the scaling complexity of stitching together air-cooled CPU servers when you need tens of thousands of cores for agent and RL environments.

Each Vera CPU is also available as a 1S or 2S server in air- or liquid-cooled form, at a 250–450 W TDP — so the same processor scales from a single node up to the full rack.

Inside the Vera CPU: 88 Olympus cores

At the heart of the platform is the Vera CPU: 88 custom Olympus cores on a single reticle-sized 3 nm compute die, with memory and I/O disaggregated into adjacent chiplets via CoWoS-R packaging. A second-generation Scalable Coherency Fabric connects all 88 cores to a shared L3 cache and memory subsystem with 3.4 TB/s of bisection bandwidth, and Spatial Multithreading creates 176 threads with partitioned core resources for predictable throughput under load.

NVIDIA Vera CPU — 88 custom Olympus cores on a reticle-sized 3 nm compute die with disaggregated memory and I/O chiplets

88 Olympus cores

Custom cores tuned for high single-thread performance — the control-heavy, latency-sensitive work behind agents and RL.

176 threads

Spatial Multithreading partitions core resources for consistent, predictable throughput rather than noisy-neighbour variance.

Up to 1.5 TB LPDDR5X

Per-CPU memory at up to 1.2 TB/s peak bandwidth, with roughly 14 GB/s provisioned per core — about 3× the per-core rate of typical data-center CPUs.

NVLink-C2C & PCIe Gen 6

1.8 TB/s NVLink-C2C, 88 PCIe Gen 6 lanes and CXL 3.1 support for tight coupling to GPUs, DPUs and storage.

The 256-CPU rack at a glance

Scaled to the full 48U rack, the platform becomes a single dense pool of CPU capacity for agent and RL environments. These are NVIDIA's stated maximum configuration figures.

256 Vera CPUs per rack
22,528 Olympus cores (45,056 threads)
400 TB LPDDR5X memory (up to)
22,500+ Concurrent CPU environments
Multiple NVIDIA Vera CPU racks deployed at scale for agentic AI sandbox environments in an AI factory
Aggregate memory bandwidth Up to 300 TB/s
L3 cache 42 GB total across the rack
Networking 64 × BlueField-4 DPUs + Spectrum-X
Form factor 48U MGX, 100% liquid-cooled

Single CPU vs full rack specs

The same Vera CPU scales from a single 1S/2S server to the 256-CPU rack. The table below summarises the per-CPU and full-rack figures.

Spec Single Vera CPU Full 48U rack
CPUs 1 Vera CPU 256 Vera CPUs
Cores / threads 88 cores / 176 threads 22,528 cores / 45,056 threads
L3 cache 164 MB 42 GB total
Memory capacity Up to 1.5 TB LPDDR5X Up to 400 TB LPDDR5X
Memory bandwidth Up to 1.2 TB/s peak Up to 300 TB/s aggregate
NVLink-C2C 1.8 TB/s 1.8 TB/s per CPU
PCIe 88 lanes Gen 6 (CPU) 22,528 lanes Gen 6
Cooling / form factor 1S/2S, air or liquid, 250–450 W TDP 48U MGX, 100% liquid-cooled

CPU sandboxes for RL and agents

The Vera CPU Rack is built around one core idea: at production scale, agentic AI and reinforcement learning need a large number of independent CPU environments running in parallel. The three patterns below are where that matters most.

Reinforcement-learning environments

RL training spins up many environments that execute actions and return rewards. Dedicated cores per environment give predictable performance instead of contention.

Agent code & tool execution

When an agent writes and runs code or calls tools, the CPU executes it in a sandbox. High single-thread performance shortens each execution and evaluation loop.

Data retrieval & processing

Agents retrieve data, run queries and process results. Large memory capacity and bandwidth keep thousands of these environments responsive at once.

How it pairs with GPU rack systems

The Vera CPU Rack is not a standalone AI trainer — it is the CPU execution layer that sits next to GPU compute. In a full AI factory, GPU racks such as the Vera Rubin NVL72 or GB300 NVL72 handle model training and inference, while the Vera CPU Rack runs the surrounding agent environments, tool calls and RL sandboxes. Both are liquid-cooled and built to share the same facility, so the liquid-cooling investment carries across CPU and GPU racks alike.

Vera CPU Rack vs GPU rack vs standard CPU servers

The specification tables above are detailed. This simpler decision table shows where the Vera CPU Rack fits compared with GPU racks and standard CPU servers, so you can see at a glance which infrastructure layer your project actually needs.

Solution Best fit When to choose it
Vera CPU Rack Agentic AI, reinforcement-learning sandboxes, CPU execution layer When your GPU infrastructure needs a dense CPU layer for sandbox, agent and environment workloads
GPU rack / GB300 / Vera Rubin Training, inference, reasoning and accelerated AI workloads When the main requirement is GPU compute for model training, inference or large AI pipelines
Standard CPU servers General enterprise, HPC and current-generation CPU workloads When you need capacity now or do not require Vera-level CPU density

Performance: what NVIDIA claims

NVIDIA positions the Vera CPU Rack as a step change for agentic CPU workloads. The figures below are NVIDIA's own stated, peak/representative claims and will vary in real deployments depending on workload, software and configuration — treat them as vendor claims pending independent benchmarks.

Up to 80% faster

Claimed workload completion versus traditional CPU infrastructure for these agentic patterns.

Up to 1.8×

Claimed acceleration over leading x86 CPUs on agentic sandbox workloads.

~3× per-core bandwidth

Up to 14 GB/s of memory bandwidth per core, roughly 3× the per-core rate of typical data-center CPUs.

3.4 TB/s fabric

Bisection bandwidth of the second-generation Scalable Coherency Fabric across the 88 cores.

Performance figures are NVIDIA's stated claims; real-world results depend on workload, software and configuration, and independent benchmarks may differ.

Before you plan a Vera CPU Rack: a 5-point checklist

Before committing to a Vera CPU Rack, work through these five questions. They drive the configuration, the cooling and networking requirements, and whether this is a deploy-now or a 2026 roadmap project — and they are exactly what our team will ask on a scoping call.

Check these 5 things first

  • How many concurrent agent or reinforcement-learning environments do you need?
  • What GPU racks will Vera support and run alongside?
  • Is your data center ready for liquid cooling?
  • What networking and DPU requirements do you have?
  • Is this a 2026 roadmap project, or do you need capacity now?

Need AI CPU capacity before Vera is available?

Vera arrives in 2026, and some workloads cannot wait. If your project needs to ship before Vera systems are commercially available, SERVER SIMPLY can help you choose current-generation AMD EPYC, Intel Xeon, GPU server or rack-scale alternatives — while keeping the Vera roadmap in mind for the next infrastructure phase.

Can't wait for Vera? Deploy now, plan the upgrade.

We'll scope a current-generation AMD EPYC, Intel Xeon or GPU rack-scale build you can deploy today, mapped to a clean migration path for when Vera ships.

Scope a deploy-now option

When a Vera CPU rack is not the right fit

The Vera CPU Rack is purpose-built for agentic AI and RL at scale. It is the wrong starting point in several common cases.

  • You need GPU training or inference: Vera is a CPU rack. For model training and inference, a GPU server or GPU rack is what you need, with Vera as a complement.
  • General-purpose or small CPU workloads: For standard enterprise or HPC CPU jobs, a conventional AMD EPYC or x86 server is more cost-effective than a dense agentic CPU rack.
  • No liquid-cooling readiness: The full rack is 100% liquid-cooled; without facility liquid cooling, a single-node Vera server or a retrofit decision comes first.
  • Workloads that ship today: Vera debuts in 2026. If you need CPU capacity now, current-generation servers are the right buy, with Vera planned for the agentic build-out.

How SERVER SIMPLY helps you deploy

Agentic AI infrastructure is a CPU-plus-GPU exercise. SERVER SIMPLY helps European teams scope and source both sides — the GPU racks that train and serve models and the CPU infrastructure that runs the agent and RL environments around them. We configure GPU servers, rack-scale systems and rack enclosures, advise on liquid-cooling readiness, and handle EU delivery, integration testing and binding lead times on quote.

What to send us for a quote

To return a specific configuration and lead-time estimate, share: your workload (agent/RL environments, training or inference) and expected scale; how many concurrent environments or GPUs you are targeting; your timeline; and your data-center power and cooling readiness, including liquid-cooling capability. We'll recommend the right balance of CPU and GPU infrastructure for your agentic AI build-out.

Get an agentic AI infrastructure configuration

Tell us about your agent and RL workloads and we'll return a CPU-plus-GPU configuration, EU delivery plan and binding lead times — for deployment now or for the next-generation build-out.

Request a configuration & lead time


FAQ

What is the NVIDIA Vera CPU Rack?

It is a dense, 100% liquid-cooled 48U rack built on the MGX platform that integrates 256 Vera CPUs and 64 BlueField-4 DPUs to run CPU sandbox environments for agentic AI and reinforcement learning at scale.

What is the Vera CPU?

A data-center CPU with 88 custom Olympus cores and 176 threads on a reticle-sized 3 nm compute die, with disaggregated memory and I/O chiplets, up to 1.5 TB of LPDDR5X and 1.8 TB/s NVLink-C2C.

How many cores are in a full Vera rack?

256 Vera CPUs give 22,528 Olympus cores and 45,056 threads, with up to 400 TB of LPDDR5X memory and up to 300 TB/s of aggregate bandwidth per rack.

What is the Vera CPU Rack used for?

Agentic AI and reinforcement learning, where models generate code and queries and CPUs execute the actions, evaluate the results and return data across thousands of parallel environments.

Does the Vera CPU Rack replace GPU servers?

No. It is the CPU execution layer that runs alongside GPU racks such as the Vera Rubin NVL72 or GB300 NVL72, which handle model training and inference.

Is the Vera CPU Rack liquid-cooled?

Yes — the full 48U rack is 100% liquid-cooled. Individual Vera CPUs are also available in 1S/2S servers in air- or liquid-cooled form at 250–450 W TDP.

How does Vera compare to x86 CPUs for agents?

NVIDIA claims up to 1.8× acceleration over leading x86 CPUs on agentic sandbox workloads and up to 80% faster completion versus traditional CPU infrastructure. These are vendor figures pending independent benchmarks.

What networking does the rack use?

64 BlueField-4 DPUs and Spectrum-X Ethernet for infrastructure services, with BlueField-4 DPU and CX9 NIC options at the server level and CXL 3.1 support.

When will the Vera CPU Rack be available?

NVIDIA has stated Vera arrives in 2026, with several major cloud and neocloud operators having committed to deploying Vera-based systems.

Can SERVER SIMPLY supply NVIDIA Vera CPU Rack systems in Europe?

SERVER SIMPLY can help European customers plan the configuration, delivery timeline and related GPU, networking, rack and liquid-cooling requirements based on vendor availability and project scope.

What should I deploy if I need agentic AI infrastructure before Vera is available?

If the project needs to go live before Vera-based systems are available, SERVER SIMPLY can help you scope current-generation CPU, GPU and rack-scale alternatives, with a migration path to Vera for the next phase.

Do I need liquid cooling before planning a Vera CPU Rack?

For the full 48U rack, yes. If the facility is not liquid-cooling ready, planning should start with power, rack, CDU and cooling readiness, or with a single-node Vera server.

Can SERVER SIMPLY help me plan agentic AI infrastructure in Europe?

Yes. We help European teams scope and source the CPU and GPU infrastructure behind agentic AI and RL — with configuration, EU delivery, integration testing, liquid-cooling planning and binding lead times on quote.