Case Study
How BenchFlow Runs High-Concurrency Agent Skill Evaluations With Daytona

15k
sandboxes provisioned in a week with Daytona
~150
engineering hours saved on sandbox infrastructure engineering
~3x
reduction in monthly sandbox infrastructure costs
BenchFlow.ai is a frontier environment lab for AI agents, backed by Google's Chief Scientist, Dropbox's Co-founder, and Founders, Inc. Its open-source research centers on three projects: SkillsBench, a benchmark for evaluating agent skill use; ClawsBench, measuring agent capability and safety in simulated workplaces; and BenchFlow, the runtime that runs those evaluations at scale.
Headquarters
San Francisco, CA
Industry
AI Research & Evaluation
Department
Engineering
Key Features
Sandbox Statefulness Infrastructure Scale Isolation and Security
Learn how this frontier agent evaluation lab partnered with Daytona to replace its self‑hosted sandbox infrastructure with a managed runtime that scales stateful agent experiments.
A single benchmark run can generate up to 15,000 independent experiments across models and tasks. That scale is only possible if you can provision thousands of isolated sandboxes at once. Daytona makes that straightforward.

Xiangyi Li
Founder of BenchFlow
01 -- CHALLENGE
Unpredictable Sandbox Demand Strained Homegrown Infrastructure
Early in his career, Xiangyi Li identified a gap in AI agent performance benchmarks: none revealed whether curated skills actually improved pass rates on real‑world tasks. To run these evaluations at scale, he built BenchFlow and SkillsBench. Delivering that capability for internal research and the broader AI community demanded sandboxed execution environments.
To keep scores reliable, BenchFlow tests 16 agent-model setups three times across 94 tasks. Because agents read files, write code, and modify filesystems as they work, each experiment needs an isolated, stateful container to prevent cross-run interference and protect the host system.
Initially, Xiangyi and his team had set up their own sandbox infrastructure using Google Cloud VMs to run Docker containers on Kubernetes. But as SkillsBench grew in scale and complexity, managing task-specific environments, lifecycle orchestration, and configuring security layers like gVisor began diverting resources from research.
The financial overhead compounded the engineering burden. Reserved VMs are billed even when idle. Supporting a thousand concurrent experiments on Google Cloud N4Ds would have cost roughly $50,000 per month, a number difficult to justify for sandbox infrastructure alone.
The complexity extended beyond BenchFlow’s internal workloads. As external events scaled up, so did the need for elastic sandbox capacity. Hackathons could involve around 30 teams running dozens of sandboxes simultaneously, while larger competitions brought more than 200 submissions. Meeting that demand on in-house infrastructure would have meant provisioning, monitoring, and tearing down thousands of sandboxes, and spending days building the access controls to protect BenchFlow’s core systems.
Those pressures led Xiangyi to search for a dedicated sandbox solution. The ideal tool would handle the underlying infrastructure out of the box, freeing BenchFlow to focus on evaluation quality. He evaluated several providers across isolation, ease of configuration, and support responsiveness, but most fell short of his requirements.
The one that didn’t was Daytona. Their agent-native runtime platform and on-demand infrastructure hit every standard he'd set.
As evaluation volume grew and we opened BenchFlow to outside teams, spending our time on the underlying infrastructure wasn’t sustainable. We needed a platform that could scale evaluations across thousands of parallel, stateful sandboxes. Daytona was the only one that could.

Xiangyi Li
Founder of BenchFlow
02 -- SOLUTION
A Managed Runtime That Delivers Thousands of Stateful, Parallel Sandboxes
Onboarding Daytona required only an API key and a single CLI parameter. After setting up both, Xiangyi had Daytona’s managed runtime platform ready to provision sandboxes at scale.
When Xiangyi runs SkillsBench today, BenchFlow automatically routes experiments through Daytona. Using Daytona’s Declarative Image Builder, his team specifies the exact base image, packages, and tools each task requires, from filesystem utilities to browser tooling. That way, each task starts from a consistent baseline without manual configuration.
Once provisioned, each sandbox enforces an isolated filesystem and process space, preventing agents from interfering with each other or destabilizing the local system. To add an extra layer of containment, Daytona enables Xiangyi and his team to create Docker-in-Docker environments by building pre-configured DinD images or manually installing Docker in a custom image.
Because Daytona handles the full sandbox lifecycle, BenchFlow achieves these outcomes with minimal lift. State persists, so agents build on prior actions the way a task requires. That consistency holds as experiment volume fluctuates. Once trials finish and results are collected, Daytona tears down sandboxes automatically, maximizing resource utilization and eliminating unnecessary spend.
With this secure, elastic infrastructure, Xiangyi and his team can trust that benchmark results accurately reflect agent performance. As a result, researchers and engineers using SkillsBench can make evidence-backed decisions about where and how to deploy agent skills.
At the same time, Daytona empowers BenchFlow to host hackathons and competitions that require flexible sandbox capacity. During these events, sandbox demand can go from zero to tens of thousands of parallel containers over a few days. Daytona absorbs those spikes without additional overhead, keeping participants’ environments fully isolated from BenchFlow's core systems.
As BenchFlow’s reach grows, the lab now hosts wider initiatives across the AI community. For the BenchFlow Agent Skill Lift competition, each submission triggers 200 Daytona sandboxes to run its evaluation, a workload that scales automatically as entries come in. Running events at that scale consistently strengthens its standing as the standard for agent skills benchmarking.
With experiments running smoothly on Daytona, Xiangyi extended its sandbox infrastructure to BenchFlow’s agent learning pipeline. As his team iterates on the SkillsBench skill packages, Daytona’s pre-configured sandboxes ensure that any performance changes result from the skills themselves.
Beyond the platform itself, Daytona’s team has been a consistent resource for Xiangyi’s team. When questions arise, responses come the same day. And when BenchFlow needs additional sandbox capacity for a benchmark run or a hackathon, Daytona adjusts resource limits on request, ensuring no experiment is left waiting.
Daytona combines on-demand infrastructure and purpose-built sandboxes to power benchmarks at any scale with minimal internal overhead. We were up and running with one API key and a single line of code.

Xiangyi Li
Founder of BenchFlow
03 -- RESULT
BenchFlow Provisions 15K Parallel Sandboxes in One Week With Daytona
With Daytona, BenchFlow can run thousands of concurrent experiments in isolated, stateful sandboxes while preserving the reproducibility that benchmark runs depend on. Xiangyi and his team now deliver reliable signals on how agent skills affect model performance with minimal infrastructure overhead. Meanwhile, hackathon participants run experiments at competition scale in fully contained environments.
15k sandboxes provisioned in a week with Daytona
~150 engineering hours saved on sandbox infrastructure engineering
~3x reduction in monthly sandbox infrastructure costs
Looking ahead, Xiangyi is planning to use Daytona’s VNC access and Web Terminal to open running sandboxes directly in the browser. This capability will enable him to inspect files and review agent trajectories significantly faster, compressing the feedback loop between runs.
Daytona’s support makes the difference. When we needed to burst capacity overnight, they adjusted our limits the same day.

Xiangyi L
Founder of BenchFlow



