Engineering case study

AI & Edge SystemsDistributed Infrastructure

ArokoDB — Embedded Vector Database for Edge AI

Search Latency (Top 10)

12 ms

p99

Active RAM Overhead

42 MB

Serialization Overhead

< 0.1 ms

Zero-copy

Works on

WindowsmacosLinuxAndroidios

Built with

RustFlutter / DartHNSW Indexmmap StorageSIMD (NEON)

01 / The Problem & Hard Constraints

The Problem & Hard Constraints

As on-device LLMs become viable on mobile hardware via LiteRT, Local RAG is the next frontier.

  • Memory Bound: 100,000 embeddings of 768 float32 dimensions consume ~300MB of raw data. Loading this into active RAM on an older iOS/Android device triggers OS-level OOM kills.
  • Latency Bound: Search queries must execute in under 16ms to maintain a 60 FPS UI experience in Flutter.

02 / Architecture & Core Design Decisions

Architecture & Core Design Decisions

  • Core Engine (Rust): Chosen for deterministic memory management and aggressive SIMD optimizations.
  • Index Structure (HNSW): Selected over IVF-Flat to guarantee logarithmic O(log N) search time for predictable low latency on mobile.
  • Storage Backend (mmap): The entire vector payload is memory-mapped to NVMe/Flash storage. The OS page cache handles swapping, keeping active RAM under 45MB.

03 / Crossing the Dart/Rust FFI Boundary

Crossing the Dart/Rust FFI Boundary

The Problem: Passing a 768-dimension float array query from Dart to Rust triggers costly serialization, spiking the Dart Garbage Collector and dropping UI frames.

The Solution: Zero-Copy Pointer Passing. The Rust core allocates the memory buffer. Dart requests a memory pointer via FFI, writes the query embedding directly into C-allocated memory, and invokes the search. Serialization overhead dropped from ~12ms to 0.04ms.

Benchmarks

Measured against the baseline

MetricStandard SQLite/JSONArokoDB (Rust + FFI + mmap)
Search Latency (Top 10)> 850 ms12 ms (p99)
Active RAM Overhead~480 MB42 MB
Serialization Overhead12 - 18 ms< 0.1 ms (Zero-copy)