Case Study
How Handshake Scales Expert-Built Reinforcement Learning Environments With Daytona

Millions
of sandboxes provisioned monthly with Daytona
1K+
Docker images maintained across workloads
120
engineering hours saved monthly

Handshake is the career network for the AI economy. It connects over 25M job seekers, 1,600 educational institutions, and one million employers, bringing career discovery, skill-building, and hiring into one place. Handshake helps people develop and demonstrate the capabilities required as the nature of work changes, while giving employers stronger signals of candidate potential. Through Handshake AI, the company also connects leading AI developers and enterprises with expert talent to help train and evaluate advanced AI systems. Learn more at joinhandshake.com.
Headquarters
San Francisco, CA
Industry
Artificial Intelligence
Department
Engineering AI Strategy and Innovation
Key Features
Sandbox Creation Speed Sandbox Statefulness Infrastructure Scale Isolation and Security Docker in Docker Container Use Daytona Volumes Sandbox Snapshots SSH Access Latency
Learn how the leading AI career network partnered with Daytona to build Reinforcement Learning Environments for frontier labs with scalable, secure sandbox infrastructure.
Daytona has been instrumental to our go-to-market motion. It enabled us to provision sandboxes instantly at scale, so we could validate and launch our reinforcement learning platform significantly faster.

Punit Arani
Senior Software Engineer at Handshake
01 -- CHALLENGE
Growing Demand for RL Environments Strained Local Sandbox Infrastructure
As frontier AI labs sought practical ways to train models on complex, real-world work, Handshake saw an opportunity to provide reinforcement learning environments (RLEs) for domain-specific tasks. To achieve that goal, Software Engineer Punit Arani and his team began building infrastructure to create, test, and refine these environments at scale. But turning that foundation into a production-grade platform required scalable, secure sandboxes.
Initially, Punit’s team used local machines to evaluate model performance across a handful of RLEs. However, as lab customers requested more diverse types of tasks, customer pilots and projects often required the team to scale from zero to thousands of concurrent sandboxes within hours and run tens to hundreds of thousands per day. At that scale, the local setup simply couldn't keep pace.
Scaling to that volume surfaced another critical requirement: enterprise-grade security and isolation. Frontier labs let models access confidential data, execute code, and build Docker containers to learn how to complete complex tasks. Each sandbox had to be fully contained, governed by strict access controls, and compliant with standards like SOC 2.
Those protections could not come at the expense of speed. Sandboxes with specialized dependencies and multi-gigabyte task data had to launch within seconds of task submission to maintain throughput. To validate each run generated reliable training data, the team also needed each sandbox to remain reproducible and available for behavior review and trajectory analysis.
Faced with these infrastructure needs, Punit began looking for a dedicated solution that could provision, isolate, and preserve sandboxes at scale. He considered building on self-managed cloud VMs and evaluated several dedicated providers. But neither path offered the combination of scale, quick time-to-value, and responsive support that would enable the team to bring Handshake’s RLE offering fully to market.
That’s when Punit discovered Daytona. Their purpose-built sandbox infrastructure was exactly what Handshake needed to deliver production-ready RLEs at scale.
Local machines couldn’t meet the demands of customer projects, especially once we needed the speed and flexibility to scale from zero to thousands of sandboxes within hours. Daytona gave us the infrastructure to meet those requirements without diverting engineering resources from building our RLE platform.

Punit Arani
Senior Software Engineer at Handshake
02 -- SOLUTION
Secure, Flexible Sandbox Infrastructure That Scales Model Runs
Daytona integrated directly into Handshake’s existing development workflow, including its GCP infrastructure and artifact registry for Docker images. In tandem, the Daytona team set up a dedicated Slack channel where Punit and his team could troubleshoot issues, share feedback, and connect directly with engineers for fast support. “We received the kind of hands-on support competitors typically reserve for enterprise customers, which was huge,” Punit shares.
With Daytona, Handshake launched and scaled its RLE offering without building and maintaining in-house sandbox infrastructure. The platform generates thousands of sandboxes on demand, enabling models to attempt expert-created tasks as soon as they’re submitted. When customer demand spikes, Punit’s team can quickly expand available capacity, scaling to tens of thousands of sandboxes within hours and sustaining that scale as needed. That flexibility enables Handshake to take on urgent, high-volume customer projects without delaying RLE delivery.
Each sandbox is fully contained, while the platform’s native Docker in Docker support enables models to build and run nested containers. Combined with enterprise-grade security (including SOC 2), these capabilities enable Punit’s team to test model execution without risking contamination between experiments or exposure of frontier labs’ proprietary data.
Each sandbox is provisioned near-instantly, even for complex, multi-gigabyte images, keeping pace with Handshake’s expert-driven workflows. This speed gives experts rapid feedback as they create and test tasks while increasing the rate at which Handshake can construct and validate new RLEs.
For repeatable environments, Punit’s team uses Daytona Snapshots to save a configured sandbox state and launch fresh instances from them on demand. As a result, they easily reproduce runs and inspect models’ trajectories to ensure each task produces reliable training data.
When an attempt needs closer investigation, developers can securely access the relevant sandbox directly through Daytona’s SSH access. That combination of reproducibility and hands-on observability helps Handshake deliver high-quality RL Environments to frontier labs.
With Daytona, we can effortlessly scale to thousands of sandboxes, and each starts fast, regardless of the image complexity. Our experts see tasks running within seconds of submitting them, which is critical for us.

Punit Arani
Senior Software Engineer at Handshake
03 -- RESULT
Handshake Provisions Millions of Sandboxes Every Month With Daytona
By partnering with Daytona, Handshake launched and scaled its RLE platform without the overhead of managing sandbox infrastructure in-house. Punit’s team can now run and analyze thousands of model attempts in parallel, delivering production-ready RLEs faster while supporting unpredictable customer demand with confidence.
Millions of sandboxes provisioned monthly with Daytona
1K+ Docker images maintained across workloads
120 engineering hours saved monthly
As more frontier AI labs adopt RL to train foundation models, Handshake will continue to expand its RLE offering. Punit is particularly excited to deepen the partnership with Daytona as new infrastructure requirements emerge, knowing the sandbox infrastructure will scale seamlessly alongside customer demand.
Daytona is at the frontier of sandbox infrastructure for AI workloads. I’m excited to scale our projects over the next year with their platform.

Punit Arani
Senior Software Engineer at Handshake



