Engineering case study

AI & Edge SystemsReal-Time Systems

CloakNotes — On-Device AI & Audio Workspace

Data Privacy

100% On-Device

Zero Cloud

Token Generation Speed

22 tokens/sec

Local NPU/GPU

Active Memory Footprint

1.1 GB

Within OS Safe Bounds

Works on

WindowsLinuxAndroidWeb

Built with

C++Dart / FlutterLiteRT (TensorFlow Lite)Whisper (Int8)INT4 / INT8 QuantizationOn-device NPU / GPU

01 / The Problem & Hard Constraints

The Problem & Hard Constraints

Modern audio workspaces lean heavily on cloud APIs (like Whisper + OpenAI) to transcribe speech and generate summaries. That convenience comes with serious tradeoffs: data privacy concerns, recurring cloud costs, and zero functionality when offline.

The real engineering challenge wasn't just "make it run locally" — it was making powerful AI fit inside the tightest resource budgets consumer devices allow, without slowing down the user experience.

  • Memory ceiling: On mobile (Android), the OS aggressively kills apps that push past ~2.5GB of active RAM. Running both Whisper and a text LLM at full size would easily need 4.8GB — over double the safe limit. This required ruthless memory optimization.
  • Real-time latency: Transcribing audio as it flows in while simultaneously generating an AI summary means the LLM needs to sustain over 15 tokens per second. That has to happen without making the device hot enough to trigger thermal throttling, or dropping UI frames.
  • Battery & heat discipline: AI inference is compute-heavy. It has to feel instant and smooth, not turn your device into a hand warmer after three minutes of use.

02 / Architecture & Core Design Decisions

Architecture & Core Design Decisions

To hit both speed and memory targets, CloakNotes splits work between two purpose-built on-device models, with all heavy computation kept off the UI thread and as close to the hardware as possible.

  • Dual-model pipeline (Whisper Int8 + SmolLM2/Phi-4 Mini): Audio gets chunked in real-time, fed into a quantized C++ Whisper engine for speech-to-text, then streamed directly into a LiteRT (TensorFlow Lite) LLM runtime to generate a clean, structured summary.
  • Native C++ engine with Dart FFI: Instead of doing tensor math in the Flutter UI isolate, every heavy operation runs in native C++ memory allocations. This keeps the UI buttery-smooth while AI inference runs at full speed.
  • Smart quantization (INT4/INT8 mixed precision): By carefully compressing model weights to a mix of INT4 and INT8 precision, total active RAM dropped from ~4.8GB down to ~1.1GB — well under the OS safety ceiling, with minimal impact to quality.
  • Hardware-aware execution: The engine routes work to the most efficient compute available (on-device NPU/GPU where supported), squeezing every watt of performance without wasting battery.

03 / Deep Technical Challenges & Solutions

Deep Technical Challenges & Solutions

The Challenge: KV-cache bloat and UI thread blocking. As recordings grow longer, the LLM's Key-Value (KV) cache expands continuously. Without a plan to manage it, this leads to out-of-memory (OOM) crashes and choppy token streaming.

The Solution: A sliding-window KV cache paired with atomic shared memory — like giving the AI a working notepad with a strict page limit, while letting the UI read results in real-time with zero copying overhead.

  1. Sliding-window KV-cache eviction (C++): Once the context exceeds ~2,048 tokens, the cache intelligently drops older, lower-attention tokens. This caps growth at a stable ~150MB, no matter how long the conversation runs.
  2. Zero-blocking streaming via atomic memory pointers: Token outputs are written directly to shared atomic memory spaces. The Dart UI polls those pointers via lightweight FFI bindings at 60 FPS — meaning zero serialization, zero data copying, and zero UI thread blocking while text streams live.
  3. Graceful memory discipline: Everything stays in tightly controlled native allocations, preventing fragmentation and keeping peak memory predictable and stable.

04 / Post-Mortem & Next Steps

Post-Mortem & Next Steps

Testing revealed thermal behavior that needed to be accounted for in sustained workloads.

  • Thermal throttling: Sustained token generation beyond ~3 minutes would heat up CPUs enough to drop generation speeds below 8 tokens/sec, hurting the real-time feel.
  • Smart thermal-aware throttling: Added dynamic batch-size scaling driven by real-time device battery and thermal telemetry. When the device starts warming up, execution intervals throttle slightly to keep temperatures steady — trading a few microseconds of speed for consistent performance and a smooth experience over long sessions.
  • Next steps: Continue tuning for newer NPUs, expand adaptive quantization profiles per device class, and refine the eviction strategy to preserve the most important context while keeping memory rock-solid.

Benchmarks

Measured against the baseline

MetricCloud API Pipeline (Whisper + Cloud LLM)CloakNotes Edge Engine
Data PrivacyTransmitted over WAN100% On-Device (Zero Cloud)
Token Generation Speed18 tokens/sec (Dependent on network)22 tokens/sec (Local NPU/GPU)
Active Memory Footprint45 MB1.1 GB (Within OS Safe Bounds)
Network Dependency100% Online100% Offline