# Contents

Agents need computers to do useful work. Those computers need filesystems, dependencies, shells, and tests to support these agents, which can mean hundreds or thousands of isolated environments at scale.

That creates a new kind of infrastructure problem.

CPUs have traditionally been judged by how quickly they finish individual jobs. Agentic workloads put them under dense, sustained pressure, with swarms of agents working concurrently as they run their own processes, tools, and tests. That makes it increasingly important for the underlying infrastructure to bring a large fleet online quickly and sustain useful work once those environments are sharing the same hardware.

To compare CPU performance on agentic workloads, we ran the same coding-agent task across large fleets of Daytona sandboxes on NVIDIA Vera and AMD EPYC Zen 5, while also measuring how quickly each system could start them.

Vera reached a peak throughput of 13.3 completed agent jobs per second and brought 1,000 sandboxes online in 10.8 seconds, giving us a concrete look at what next-generation CPU performance can mean for dense agent infrastructure.

Bringing Thousands Of Agent Environments Online

Coding agents proved that getting work done requires a shell, a filesystem, and a development environment. As agents evolve into more capable personal assistants and general-purpose tools, that foundational toolkit needs to expand with them.

progression of llms
Progression of LLMs

Each agent needs an environment with its own files, tools, and dependencies. If every environment required a dedicated physical machine, supporting agent workloads at that scale would mean buying and managing a machine for every task. That would be wasteful and difficult to scale, so we need to fit hundreds or thousands of isolated environments onto the machines we already have.

That led us to the first thing we wanted to measure with Vera

How quickly can one CPU bring a large group of agent environments online?

We started with 1,000 sandboxes. Each one booted the same agent image and ran the same commands. We measured the time it took to provision all the sandboxes and get each one ready to accept work.

On NVIDIA Vera, all 1,000 sandboxes were ready in 10.8 seconds. The EPYC 9575F took 32.3 seconds, while the EPYC 9755 took 88.4 seconds.

vera graph showing create and ready for 1k sandbox
Vera created 1,000 sandboxes in just 10.8s

We doubled the fleet to 2,000 sandboxes to see how Vera handled an even larger burst.

Vera brought all 2,000 sandboxes online in 27.4 seconds.

“It’s like a Rimac Nevera hitting 200 km/h before a Porsche Carrera clears 100.”

— Ivan Burazin, Co-founder & CEO

This matters because Daytona is constantly creating new environments. On a typical day, Daytona spins up millions of sandboxes, and the median sandbox lives for less than five minutes. Agents finish tasks, new ones start, rollouts reset, and environments get replaced. At that scale, startup is not something that happens once at the beginning of a workload. It is happening all day.

Vera handled that burst well. It could bring 2,000 isolated environments online before the 9575F finished bringing 1,000 online.

As important as startup is, it is only the beginning of an agent’s lifecycle.

Once all the environments are running, they begin competing for the same underlying machine resources, from CPU time and cache to memory bandwidth, RAM, and I/O.

A New Kind of Compute Density

Software infrastructure was built around people using computers. A company with 1,000 engineers might have 1,000 laptops, but there was no reason to pack all 1,000 development environments onto a single CPU. Agents change that assumption. A single organization may eventually run more agents in parallel than it has engineers, each needing its own isolated filesystem, processes, dependencies, and tools.

We need to fit far more computers onto each physical machine than we ever needed to for humans

It is not enough to make one sandbox fast. The underlying system has to start many sandboxes quickly, isolate each one safely, and continue doing useful work as they compete for the same CPU.

How Much Work Can One Machine Sustain?

Next, we moved from measuring startup to testing sustained throughput under load. We increased concurrency from 44 to 2,000 sandboxes, each running the same coding-agent workload.

  • Build a small broken python package

  • Search through it

  • Apply a patch

  • Run tests

As we increased the load, the three CPUs began to separate in how much work they could sustain.

agent throughput vs concurrency
Throughput from 44 to 2,000 concurrent sandboxes
  • EPYC 9575F: levels off around 8 agent jobs/sec

  • EPYC 9755: levels off around 11 agent jobs/sec

  • NVIDIA Vera: reaches 13.3 agent jobs/sec and stays close to 13 through 2,000 sandboxes

Even at 2,000 concurrent sandboxes, Vera is still completing 12.7 agent jobs/sec, compared with 10.9 on the 9755 and 8.1 on the 9575F.

Vera sustained high throughput even as sandbox concurrency exceeded the socket’s capacity. Once the socket is full, total throughput stays close to 13 completed agent jobs per second all the way out to 2,000 sandboxes. The machine keeps producing agent work at nearly the same rate.

packed tput
Peak throughput for each chip

That is the part that matters for dense agent workloads. You can keep adding isolated environments without total useful work falling apart.

How Density Translates To Cost

Throughput per machine tells us how much hardware we need to do the same amount of agent work.

Take an illustrative workload that needs to sustain 1,000 completed agent jobs per second.

Under that model, Vera needs about 15% fewer machines than the 9755 and about 38% fewer than the 9575F.

machines need
Machines needed for 1,000 jobs/sec

That difference shows up beyond the CPU itself. Fewer machines means fewer servers to provision, less rack space, fewer network connections, less power and cooling, and fewer servers for Daytona to operate.

“Same budget, more than 50% more agent work. For our customers, that's the whole story.”

— Kevin Qi, Head of Finance

What This Means for Daytona Customers

This is a win for both Daytona and the teams building on top of it.

Vera gives Daytona better infrastructure for agent workloads. We can bring environments online faster, pack more of them onto a machine, and sustain more useful agent work once the machine is full. Customers get those gains through Daytona without having to manage the hardware underneath.

Handshake offers a concrete example. It builds reinforcement learning environments where AI models learn by attempting tasks and evaluating their results. When a lab starts a project, the workload can jump from zero to thousands of concurrent sandboxes within hours, with daily sandbox runs reaching six figures. Each attempt needs its own isolated environment, so the number of experiments a team can run depends in part on how much useful work each machine can sustain.

Reducing infrastructure bottlenecks lets researchers run more experiments, complete more agent work, and iterate toward increasingly capable systems faster.

Preparing Infrastructure for Agent Scale

Agents are changing what we ask from infrastructure. Instead of giving one person one computer, we are starting to pack hundreds or thousands of isolated computers onto the same physical machine.

That changes what good performance looks like. We care about how quickly the environments can be provisioned, how much useful work the machine can sustain once it is full, and how efficiently we can turn hardware into completed agent work.

Vera performed well on both sides of that problem. Daytona brought 1,000 sandboxes online in 10.8 seconds, pushed that to 2,000 in 27.4 seconds, and still sustained close to 13 agent jobs per second at the highest levels of concurrency we tested.

The bigger takeaway is that agent infrastructure is becoming a density problem. As agent fleets grow, the systems underneath them need to create more environments, pack them more efficiently, and keep them productive under heavy load.

“The number that interests me isn't the peak. It's what happens after the machine is full. Vera keeps running agent workloads at almost the same rate all the way to 2,000 sandboxes. It slows down a little but stays stable instead of falling apart. Thousands of agents on a single machine, and Vera keeps them alive. That's what matters.”

— Vedran Jukić, Co-founder & CTO

Daytona handles this infrastructure complexity so that developers can build agents (and agents can build agents).