Engineering case study

Distributed InfrastructureSecurity & Networking

Eruption — Low-Latency API Gateway & Bot Shield

P99 Proxy Latency Overhead

1.8 ms

Max Sustainable RPS (Single Node)

28,500 RPS

Memory Footprint at 10k Connections

85 MB

Works on

LinuxDocker

Built with

RustTokioRedisLuaZero-Copy TCP Proxy

01 / The Problem & Hard Constraints

The Problem & Hard Constraints

Standard reverse proxies like NGINX are exceptional for static routing, but adding dynamic bot mitigation, payload inspection, and distributed rate-limiting often introduces unacceptable latency.

  • Latency Overhead Budget: The gateway must inspect the payload, check the rate limit, and proxy the request adding no more than 2ms of overhead at the 99th percentile (p99).
  • High Concurrency (C10K+): The system must sustain 25,000+ concurrent TCP connections without memory bloat or thread starvation.
  • Distributed State: Rate limits must be synchronized across multiple edge nodes globally without synchronous database locks blocking the main request thread.

02 / Architecture & Core Design Decisions

Architecture & Core Design Decisions

  • Core Engine (Rust / Tokio): Built on Rust’s asynchronous runtime (tokio) to utilize a thread-per-core, event-driven architecture, avoiding context-switching overhead under heavy concurrent load.
  • Zero-Copy Proxying: Implemented zero-copy memory boundaries. Incoming TCP stream payloads are evaluated in place and passed to the upstream server without allocating new memory buffers, heavily reducing Garbage Collection/Drop overhead.
  • Multi-Tiered Rate Limiter: Instead of making a synchronous Redis call for every request, Eruption uses an L1/L2 caching strategy.
  • L1 (Local): In-memory Token Bucket using lock-free atomic counters for instant evaluation.
  • L2 (Global): Asynchronous batched pipelines sync the local counters to a central Redis cluster every 500ms using Lua scripts to maintain global state.

03 / Deep Technical Challenges & Solutions

Deep Technical Challenges & Solutions

The Challenge: Thread Blocking on Malicious Payload Inspection. Running Regex-based WAF (Web Application Firewall) rules on incoming JSON payloads was causing CPU spikes, blocking the async event loop, and starving healthy connections.

The Solution: Bounded Execution and Streaming Parsers. Instead of loading the entire payload into memory to run Regex, I implemented a streaming JSON state machine. The shield inspects the byte stream in chunks as it arrives over the socket. If a malicious signature (e.g., SQL injection vectors or buffer overflow padding) is detected in the stream, the socket is immediately terminated (TCP RST) before the rest of the payload is even downloaded, saving massive CPU cycles and memory bandwidth.

04 / Post-Mortem & Next Steps

Post-Mortem & Next Steps

  • TCP Exhaustion: During stress testing at 30k RPS, the host OS ran out of ephemeral ports, causing socket binding failures.
  • Mitigation: Tuned the Linux kernel sysctl parameters (tcp_tw_reuse, ip_local_port_range) and implemented aggressive TCP Keep-Alive pooling for upstream connections to prevent connection churn.

Benchmarks

Measured against the baseline

Simulated load-test metrics using wrk2 on a 4-core, 8GB instance.

MetricStandard API Gateway (e.g., Kong DB-less)Eruption (Rust + Async Sync)
P99 Proxy Latency Overhead4.5 ms1.8 ms
Max Sustainable RPS (Single Node)~14,000 RPS28,500 RPS
Memory Footprint at 10k Connections~350 MB85 MB
Rate-Limit Evaluation Time1.2 ms (Redis Block)0.05 ms (L1 Atomic Cache)