Engineering case study

Distributed InfrastructureAI & Edge Systems

Stralis — Distributed Media Automation Engine

Max Task Duration

Unbounded

Hours/Days

Failure Blast Radius

Isolated to Single Node

Broker Memory / 1k Jobs

12 MB

Claim-Check Pattern

Works on

LinuxWebDocker

01 / The Problem & Hard Constraints

The Problem & Hard Constraints

Automating heavy media processing (e.g., 4K video transcoding, multi-gigabyte audio transcription via Whisper) fundamentally breaks standard synchronous web architectures.

  • HTTP Timeout Wall: Most API gateways (like AWS API Gateway or Cloudflare) forcefully close connections after 30–60 seconds. Video/AI processing can take hours.
  • State Machine Failures: In a 5-step visual node pipeline, if the GPU worker crashes on Step 4 due to an Out-Of-Memory (OOM) error, the user cannot afford to re-run Steps 1 through 3.
  • Message Queue Bloat: You cannot pass 2GB video files inside JSON payloads through Redis or RabbitMQ without crashing the broker.

02 / Architecture & Core Design Decisions

Architecture & Core Design Decisions

  • Directed Acyclic Graphs (DAGs): Visual pipelines are compiled into mathematical DAGs. The orchestration engine traverses the graph, ensuring dependent nodes only execute when parent nodes successfully resolve.
  • Asynchronous Worker Pools: The API gateway immediately returns an HTTP 202 (Accepted) with a job_id. Background worker nodes poll a distributed message broker (Redis Streams/RabbitMQ) to claim and process jobs independently.
  • Claim-Check Messaging Pattern: Instead of passing media through the message queue, workers upload the intermediate file to an S3-compatible Object Store. Only the lightweight URI/reference (the “claim check”) is passed via the message queue to the next node.

03 / Deep Technical Challenges & Solutions

Deep Technical Challenges & Solutions

The Challenge: Handling GPU Worker Crashes & Pipeline Resumption. AI inference nodes frequently crash due to hardware faults or VRAM exhaustion. Standard systems fail the entire pipeline, wasting compute resources and time.

The Solution: Checkpointing & Idempotent Executions. Every node execution was made strictly idempotent. Before executing a node, the worker queries the database to check if a valid output artifact already exists for that specific node’s hash. If a worker dies mid-execution, the orchestrator detects the missed heartbeat and requeues the task. The new worker resumes the DAG exactly at the failed node, utilizing the intermediate S3 artifacts generated by previous steps.

Benchmarks

Measured against the baseline

MetricSynchronous API ArchitectureStralis Distributed DAG Engine
Max Task Duration< 60 seconds (Gateway Timeout)Unbounded (Hours/Days)
Failure Blast RadiusEntire Pipeline FailsIsolated to Single Node
Broker Memory / 1k Jobs~2.5 TB (Payload in Queue)12 MB (Claim-Check Pattern)
Pipeline Resumption Time100% Restart Penalty0% (Resumes at Failure Point)