Engineering case study

Security & SystemsDistributed Infrastructure

Crucible — Secure High-Speed Execution Sandbox

Cold-Start Startup Time

4.2 ms

Total Round-Trip Latency

14 ms

p99

Isolation Boundary

Hardware-Level KVM Hypervisor

Works on

Linux

Built with

Firecracker Micro-VMLinux KVMcgroups v2Seccomp BPFUnix Domain Sockets

01 / The Problem & Hard Constraints

The Problem & Hard Constraints

Running untrusted, arbitrary user-submitted code (e.g., LeetCode-style execution platforms) poses severe security and system availability risks.

  • Container Breakout & Isolation: Preventing malicious syscalls (e.g., unauthorized network access, fork bombs, file system tampering) from compromising host infrastructure.
  • Cold-Start Latency Budget: Code execution environments must spin up, execute, capture stdout/stderr, and tear down in under 20ms to provide an instant developer response.
  • Strict Resource Limits: Each runtime instance must be capped at 64MB RAM and 0.5 vCPU cores without kernel panic or leaking memory state back to the host.

02 / Architecture & Core Design Decisions

Architecture & Core Design Decisions

  • Micro-VM Isolation (Firecracker / Linux KVM): Selected micro-VMs over standard Docker containers. Docker shares the host kernel, leaving a surface area for kernel exploits. Micro-VMs provide hardware-level isolation with minimal overhead.
  • Unix Domain Socket (UDS) Agent: Communication between the API gateway and worker execution nodes bypasses TCP/IP network overhead, transmitting raw binary payloads over local Unix Domain Sockets for minimal IPC latency.
  • cgroups v2 + Seccomp BPF Filtering: Strict Linux kernel syscall restrictions enforce a zero-trust boundary, blocking forbidden syscalls (e.g., socket creation, process cloning) at the kernel boundary before execution starts.

03 / Deep Technical Challenges & Solutions

Deep Technical Challenges & Solutions

The Challenge: Micro-VM Cold Starts. Spawning a fresh micro-VM per execution takes ~100ms—far too slow for real-time code evaluation.

The Solution: Pre-Warmed Micro-VM Snapshot Pools. I architected a background worker pool that maintains “paused” micro-VM snapshots in memory.

Result: Reduced total execution round-trip latency from 180ms to 14ms (p99).

  1. An active snapshot is restored in <4ms via memory page-mapping.
  2. The user code is injected directly into a dedicated RAM region over UDS.
  3. Execution completes, metrics are logged, and the VM instance is destroyed.

04 / Post-Mortem & Next Steps

Post-Mortem & Next Steps

  • Memory Pressure Under Spikes: Initial testing revealed memory exhaustion under heavy burst traffic due to slow garbage collection of destroyed VM instances.
  • Mitigation: Implemented a fixed-size ring buffer for VM worker slots, applying backpressure at the gateway layer when all worker slots are saturated.

Benchmarks

Measured against the baseline

MetricStandard Docker ExecutionCrucible Micro-VM Sandbox
Cold-Start Startup Time120 ms - 250 ms4.2 ms
Total Round-Trip Latency180 ms14 ms (p99)
Isolation BoundaryShared OS KernelHardware-Level KVM Hypervisor
Memory Footprint / Worker180 MB38 MB